MySQL JDBC Configuration for High-Performance Batch Jobs

Connector/J options that matter for batch workloads on MySQL - useServerPrepStmts, cachePrepStmts, rewriteBatchedStatements, result-set streaming with Integer.MIN_VALUE, and server-side cursors with useCursorFetch - with notes on how each interacts with Spring Batch's JdbcCursorItemReader.

Batch jobs stress a JDBC driver differently from a web application. They run the same statement many times, insert or update rows in bulk, and read result sets that do not fit in memory. MySQL Connector/J has several connection options that change how it behaves in each of those situations, and most of them are off by default. This article walks through the options that matter most for batch work, what each one actually does on the wire, and the pitfalls to watch for when combining them.

Using Server-Side Prepared Statements

PreparedStatement helps optimize repeated query execution. Instead of declaring the entire SQL string such as SELECT * FROM CITY WHERE COUNTRY = 'KOREA' AND POPULATION > 10000, it separates the static SQL structure SELECT * FROM CITY WHERE COUNTRY = ? AND POPULATION > ? from the dynamic parameters.

When a query is executed repeatedly, the server reuses the parse tree built by parsing the SQL and executes it with different parameters. This reduces CPU and other resource usage on the server. If the JDBC driver sends only the varying parameters instead of transmitting the entire query for each execution, the amount of network traffic can also be reduced. Whether such optimizations actually occur depends on the DBMS and configuration options.

MySQL provides the useServerPrepStmts option to control whether PreparedStatement optimization is performed on the server side.

Since the default value is false, the JDBC driver sends the fully substituted SQL string to the server for every execution when no additional configuration is applied. This is referred to as client-side PreparedStatement or emulated PreparedStatement. This default behavior has the advantage of requiring only one network round-trip per query, but server-side optimizations are not applied.

In environments such as typical web applications, where a variety of queries are executed, it may be more efficient to keep the default behavior while enabling cachePrepStmts=true and increasing related cache sizes. [1]

Setting useServerPrepStmts=true enables PreparedStatement parsing on the server side. When a PreparedStatement object is created, the template query containing ? placeholders is first sent to the server, parsed, and prepared. Each time execute() is called, only the parameter values are transmitted, and the already-prepared query is executed. If the same query is executed repeatedly, the parse tree built from SQL parsing is reused, improving performance and reducing network bandwidth. However, MySQL clears the optimization state once execution finishes and optimizes again on the next execution, so you cannot assume the execution plan is built only once and reused indefinitely. [2] On the other hand, since at least two network round-trips (prepare, execute) are required to execute a query, performance may actually degrade for queries that are executed only once.

If your application repeatedly executes the same query in MySQL, the useServerPrepStmts=true option is worth trying out. However, MySQL parses quickly and rebuilds the execution plan on every execution, so the gain from this option may be small. In the two performance tests introduced in the footnote above, server-side prepared statements were not faster than client-side ones. When enabling this option, also enable cachePrepStmts=true so that the prepare request does not add a network round-trip on every execution, and measure performance with your actual batch job to decide whether to keep it.

In MySQL versions prior to 5.1.17, using this option prevented queries from using the query cache, but from 5.1.17 onwards it can be used together with the query cache. Note that the query cache itself was deprecated in MySQL 5.7.20 and removed in 8.0, so this constraint no longer matters on 8.0 and later. [3] In MariaDB 10.6 and later, additional optimizations reduce metadata retransmission when using useServerPrepStmts=true, resulting in even better performance. [4]

Improving Batch Update Performance

When inserting or updating multiple rows, using the JDBC Statement.executeBatch() method can execute them faster than repeatedly calling executeUpdate() for single statements. MySQL can further optimize batch update performance by enabling the rewriteBatchedStatements=true option in the JDBC connection URL.

Since the default value is false, it must be explicitly enabled. When enabled, the MySQL JDBC driver combines multiple individual queries into a single statement. For example, let’s look at inserting three rows with a batch update using the following INSERT query:

Single INSERT form
INSERT INTO access_log(access_date_time, ip, username) VALUES (?, ?, ?);

Without rewriteBatchedStatements=true, the driver executes the above statement three times. With the option enabled, the driver merges them into a single statement:

Merged INSERT form
INSERT INTO access_log(access_date_time, ip, username) VALUES
  (?, ?, ?),
  (?, ?, ?),
  (?, ?, ?);

Since the server processes the merged INSERT as a single statement, one execution of that statement takes longer. In replicated environments, this can increase the burden of replication lag.

Even when multi-value VALUES clauses are not possible, the driver still combines multiple INSERT or UPDATE statements separated by ; into a single transmission when the batch contains four or more statements. Batches of three or fewer statements are executed one by one. The direct benefit of this mode is packing multiple SQL statements into one request to reduce network round-trips. Unlike the multi-value INSERT, which the server processes as a single statement, the ;-combined form merely transmits several statements at once, so the parsing and execution of each statement are not merged into one.

Be aware that rewriteBatchedStatements=true may conflict with other JDBC options or require additional tuning. The Connector/J documentation once stated that rewriteBatchedStatements was ignored when used together with useServerPrepStmts=true, or with useCursorFetch=true (introduced later in this document) which implicitly enables it. The 8.0.30 release notes corrected this as a documentation error. [5] Query rewriting now applies as-is even when combined with server-side PreparedStatements.

When combining queries, the driver calculates on its own how many statements to merge at once so that the combined statement stays within the max_allowed_packet limit. Therefore a small value usually does not cause an error; it just reduces how many statements are merged at once, shrinking the benefit of combining queries. To insert or update large volumes at once, it is better to set this value generously.

max_allowed_packet can be set in the MySQL configuration file or as a global system variable. The session value is read-only and initialized from the global value when the connection is established, so if you change the global value dynamically, new connections use the changed value. [6] You can check it with: SHOW VARIABLES LIKE 'max%'; Note that the bulk_insert_buffer_size setting is only referenced by the MyISAM engine—which is rarely used today—and can be ignored when using InnoDB.

Due to such conflicts and interactions among options, it may be difficult to find a single DB configuration optimized for both reads and writes. Another possible approach is to declare separate data sources in the application for large-scale read and write operations.

In Connector/J versions up to 8.0.28, inserting BLOB (Binary Large Object) values with a batch update could cause a NullPointerException, but this has been fixed in newer versions. [7]

Options for Large-Scale Data Retrieval

If an application using Spring Batch with MySQL executes queries through JdbcCursorItemReader with default settings, the entire result set will be fetched at once and loaded into the application’s memory. When querying large datasets, this can cause Out Of Memory (OOM) errors. To avoid this, you must use ResultSet streaming or server-side cursors.

Streaming Results One Row at a Time

ResultSet streaming retrieves query results gradually instead of receiving them all at once.

To use this approach, configure the PreparedStatement as follows:

Creating PreparedStatement for streaming
PreparedStatement statement = con.prepareStatement(
    sql,
    ResultSet.TYPE_FORWARD_ONLY,
    ResultSet.CONCUR_READ_ONLY
);

statement.setFetchSize(Integer.MIN_VALUE);

Note that calling setFetchSize(Integer.MIN_VALUE) is not technically valid per JDBC specification. The Javadoc for the java.sql.Statement interface states that if a value less than 0 is passed to the setFetchSize(int) method, it should throw a SQLException. However, since MySQL Connector/J enables streaming mode by passing Integer.MIN_VALUE to this method, developers have no choice but to use it that way.

While the streaming mode reduces memory usage, it is not always advantageous. The query request is sent only once and the server pushes the entire result back as consecutive packets, so it does not incur one network round-trip per row. [8] Instead, while the driver reads and processes the result one row at a time, server resources and locks are held until the query completes, so the slower the application consumes the rows, the longer this burden lasts. Other drawbacks are that you cannot directly control the size of each transfer, and no other queries can be executed on the same connection until the ResultSet is closed. [9]

In Spring Batch, calling JdbcCursorItemReader.setFetchSize(Integer.MIN_VALUE) enables the streaming mode. Statement creation options such as ResultSet.TYPE_FORWARD_ONLY are applied internally within JdbcCursorItemReader. If the JdbcCursorItemReader.verifyCursorPosition property remains at its default true, it conflicts with TYPE_FORWARD_ONLY and produces the following error:

Error caused by verifyCursorPosition=true
org.springframework.dao.TransientDataAccessResourceException: Attempt to process next row failed; SQL [SELECT * FROM access_log]; Operation not allowed for a result set of type ResultSet.TYPE_FORWARD_ONLY.

Accordingly, a JdbcCursorItemReader configured for per-row streaming should be created as follows:

JdbcCursorItemReader for streaming queries in MySQL
return new JdbcCursorItemReaderBuilder<T>()
  .name("streamingDbReader")
  .dataSource(this.dataSource)
  .sql(sql)
  .rowMapper(rowMapper)
  .fetchSize(Integer.MIN_VALUE)
  .verifyCursorPosition(false)
  .build();

Using Server-side Cursors

MySQL server-side cursors work by storing results in a temporary table and allowing the client to fetch the data in configurable chunks. Supported in server versions newer than MySQL 5.0.2, they can be enabled by adding useCursorFetch=true to the JDBC URL. [10] Since the default value is false, this option must be explicitly enabled if client-side cursors are not desired.

Even with useCursorFetch=true, server-side cursors will not be used unless a positive fetch size is specified. This can be configured per query using the Statement.setFetchSize(int) method in the JDBC API, or a default value can be set via the defaultFetchSize property in the JDBC connection URL. [11]

In Spring Batch, the fetch size can be specified using JdbcCursorItemReader.setFetchSize(int). It is recommended to set this value equal to chunk size in your step configuration.

As mentioned earlier, enabling the useCursorFetch=true option also automatically enables useServerPrepStmts=true.


1. For performance test cases combining useServerPrepStmts and cachePrepStmts, see the following articles: https://vladmihalcea.com/mysql-jdbc-statement-caching/ , https://tech.kakaopay.com/post/how-preparedstatement-works-in-our-apps/
2. The MySQL 8.4 manual describes what is cached for prepared statements as an internal structure converted from the SQL; for example, SELECT * is stored expanded into the actual column list. https://dev.mysql.com/doc/refman/8.4/en/statement-caching.html Query_expression.clear_execution() in the MySQL 8.4 source also resets the optimized state to false before a prepared statement is re-executed. https://dev.mysql.com/doc/dev/mysql-server/8.4.9/sql__lex_8h_source.html
3. The deprecation and removal of the query cache is documented in the 'How the Query Cache Operates' section of the MySQL 5.7 Reference Manual: https://dev.mysql.com/doc/refman/5.7/en/query-cache-operation.html
4. This improvement is tracked in https://jira.mariadb.org/browse/MDEV-19237
5. The correction can be found in the following release notes: https://dev.mysql.com/doc/relnotes/connector-j/en/news-8-0-30.html
6. The max_allowed_packet entry in the MySQL 8.4 Reference Manual states that the global value can be changed dynamically but the session value is read-only. https://dev.mysql.com/doc/refman/8.4/en/server-system-variables.html#sysvar_max_allowed_packet
8. The response structure, in which the server answers a single query request with one packet per row sent in sequence, is described in the following MySQL protocol documentation: https://dev.mysql.com/doc/dev/mysql-server/latest/page_protocol_com_query_response_text_resultset.html
9. This is explained in the ResultSet section of the following Connector/J implementation notes. It describes the restriction that you must read all of the rows (or close the ResultSet) before issuing any other queries on the same connection, and introduces the useCursorFetch option as an alternative that fetches a set number of rows at a time. https://dev.mysql.com/doc/connector-j/en/connector-j-reference-implementation-notes.html
10. The Connector/J documentation used to describe this property as determining whether cursor-based fetching should be used when the server version is newer than MySQL 5.0.2 and the fetch size is greater than 0. As the minimum supported server versions moved up, this version condition was dropped from recent documentation: https://dev.mysql.com/doc/connector-j/en/connector-j-connp-props-performance-extensions.html

Running External Processes from Java: JDK 25 on Linux 6.x

What to get right when launching external processes from Java on JDK 25 and Linux 6.x: draining stdout and stderr, closing stdin, timeouts, termination policy, and cleaning up descendants with process groups or cgroups. Compares zt-exec and Apache Commons Exec, and explains the default posix_spawn launch mechanism, the cost of FORK on large heaps, and the removal of VFORK.

This article, based on JDK 25 and Linux 6.x, covers what to watch for when writing Java code that launches external processes, and the libraries you can use for it. It then looks at how the JVM creates child processes on Linux: the posix_spawn() function, the behavior of jspawnhelper, a small executable shipped with the JDK, the launch latency that appears with large heaps, and the setting that changes the launch mechanism. The results in this article were verified in the environment below. The example code is in the GitHub repository.

Item Value

OS

Ubuntu 24.04.4 LTS, kernel 6.17

Architecture

x86-64

glibc

2.39

JDK

Temurin 25+36

zt-exec

1.13.0

Apache Commons Exec

1.6.0

Pitfalls of Launching External Processes with the JDK Classes

ProcessBuilder vs. Runtime.exec

An external process launched from Java code is handled through a java.lang.Process object. You get a Process object from ProcessBuilder.start() or Runtime.exec(). ProcessBuilder was added in JDK 5.0 (Java 1.5), and from that release the Sun JDK’s Runtime.exec() was also changed to delegate internally to ProcessBuilder. A contemporary write-up that examined the JDK 1.5.0-beta2 source shows this implementation as well. The Runtime source in OpenJDK 25 likewise creates a ProcessBuilder and calls start(). The two APIs therefore share the process creation implementation; the main difference is how they handle the launch configuration.

I recommend ProcessBuilder for new code. It offers a more convenient API for specifying environment variables, the working directory, and output redirection. The output pipe handling discussed in this article can also be configured with redirectOutput(), redirectError(), redirectErrorStream(), and similar methods.

Item Runtime.exec() ProcessBuilder

When it runs

Runs when exec() is called

Runs on start() after the configuration is set up

Command and arguments

A single string or an array of strings

Varargs or a list with the arguments already separated

Environment variables

Passed as an array of NAME=value strings

Modified through the Map returned by environment()

Working directory

Passed as an argument to an overload

Set with directory()

I/O configuration

No API for redirection or for merging standard error

Provides redirectOutput(), redirectError(), redirectErrorStream(), inheritIO(), and more

In particular, the Runtime.exec(String) family, which takes the command as a single string, has been deprecated since JDK 18. The Runtime Javadoc explains that these methods split arguments on whitespace alone, so they can mishandle things like file names that contain spaces. Adding quotes does not group arguments the way a shell would.

// Cannot pass a file name with a space as a single argument
Runtime.getRuntime().exec("cat \"my file.txt\"");

// Both forms pass the file name as a single argument
Runtime.getRuntime().exec(new String[]{"cat", "my file.txt"});
new ProcessBuilder("cat", "my file.txt").start();

Limits of the ProcessBuilder Defaults

In Java, you can launch an external process with a few simple lines of code using the ProcessBuilder class, as shown below.

Calling ProcessBuilder with the defaults
Process process = new ProcessBuilder("echo", "hello").start();
int exitCode = process.waitFor();
System.out.println("exit=" + exitCode);

With echo hello, which produces little output and finishes quickly, this code works without problems. But if you run a command that produces a lot of output, waits on standard input, or never finishes, this code can hang.

Handling the stdout and stderr Pipes

If the child process fills the limited capacity of the standard output (stdout) or standard error (stderr) pipe while the parent process is not reading, the child blocks on its next write. To prevent that, you need code that keeps consuming standard output and standard error. You must use one of the following two approaches.

  • Read the standard output and standard error streams on separate threads.

    • If you need to process or refer to the two outputs separately, read each on its own thread.

    • If you do not need to tell the two outputs apart, you can merge them into one and read it on a single thread.

  • If you do not need to read a stream at all, redirect it to a file, to the parent’s output, or to /dev/null.

The earlier 'Calling ProcessBuilder with the defaults' code used neither of these approaches. The following sections reproduce how this code hangs with examples and explain how to fix it.

Deadlock from an Unread Pipe Buffer

Even with ProcessBuilder, the child process’s standard output and standard error are delivered to the parent process through separate pipes by default. If the parent does not read a pipe, the pipe buffer fills up and the child process cannot finish writing.

This problem occurs even if you never call waitFor(). The parent can go on executing the code that follows the child’s launch, but the child may be unable to finish its work because its write() call to the full pipe blocks. If on top of that the parent just waits for the child to exit with waitFor(), the two end up in a deadlock, each waiting for the other. So removing waitFor() does not solve it; you have to read or redirect the output pipes. The Process Javadoc in JDK 25 warns about this blocking and the possibility of deadlock.

From the Javadoc of the Process class

Because some native platforms only provide limited buffer size for standard input and output streams, failure to promptly write the input stream or read the output stream of the process may cause the process to block, or even deadlock.

The pipe(7) manual page tells how large the "limited buffer size" mentioned in the Javadoc is on Linux. Since kernel 2.6.11 the default pipe capacity has been 16 pages, that is, 64KiB with 4KiB pages. However, it can be smaller depending on the per-user pipe memory limit or system settings, and it can be changed with F_SETPIPE_SZ, so you should not assume 64KiB as a fixed limit.

The code below has the seq 1 100000 command print 588,895 bytes and waits 3 seconds without reading them.

DeadlockDemo.java
Process process = new ProcessBuilder("seq", "1", "100000").start();
boolean finished = process.waitFor(3, TimeUnit.SECONDS);
System.out.println("finished within 3s: " + finished + ", alive: " + process.isAlive());
long lines = process.inputReader().lines().count();
int exitCode = process.waitFor();
System.out.println("read " + lines + " lines, exit=" + exitCode + ", alive: " + process.isAlive());
Output of DeadlockDemo.java
finished within 3s: false, alive: true
read 100000 lines, exit=0, alive: false

After the parent process has waited 3 seconds, the child process is still alive. Once the pipe buffer filled up, the write() call in seq blocked and never returned. When the parent read the pipe to the end, the child wrote the rest of its output and finished with exit code 0 (alive: false). Without the timeout, the parent would have kept waiting in waitFor() as well.

The inputReader() method used in the example was added in JDK 17. Before that, you had to wrap getInputStream() in an InputStreamReader and a BufferedReader.

Using inputReader()
// Before JDK 17
BufferedReader reader = new BufferedReader(
        new InputStreamReader(process.getInputStream()));
long lines = reader.lines().count();

// JDK 17 and later
long lines = process.inputReader().lines().count();

The two snippets choose their default encoding differently. inputReader() decodes with the encoding named by the native.encoding system property, while an InputStreamReader created without arguments uses Charset.defaultCharset(). Since JDK 18 the default charset is UTF-8, but if you specify -Dfile.encoding=COMPAT, the default encoding is determined by the operating system, the locale, and so on, as it was before.

inputReader() does not detect the child process’s output encoding automatically. If the child writes in the same encoding as the parent JVM’s native.encoding, this default is appropriate. If the child explicitly writes UTF-8 or runs under a different locale, you must specify the encoding yourself to match its output, for example inputReader(StandardCharsets.UTF_8).

Reading Both Streams Concurrently

If you need to receive and process standard output and standard error separately, you have to drain both pipes at the same time. Reading one pipe to the end and then reading the other is not enough. If the standard error pipe fills first, the child process cannot write any more standard output, and the parent deadlocks waiting for EOF on standard output. In this case the program hangs in the read before it even reaches waitFor().

The code below runs a command that writes 588,895 bytes to standard error and then one line to standard output, and it reads standard output to the end before reading standard error.

SequentialReadDemo.java
Process process = new ProcessBuilder("sh", "-c", "seq 1 100000 >&2; echo done").start();
System.out.println("reading stdout...");
long outLines = process.inputReader().lines().count();
long errLines = process.errorReader().lines().count();
System.out.println("stdout " + outLines + ", stderr " + errLines + ", exit=" + process.waitFor());

When you run this code, nothing follows the first line no matter how long you wait, and ps shows that the child seq process is still alive.

Output of SequentialReadDemo.java
reading stdout...

seq is blocked because the standard error pipe is full, and the parent is waiting for an EOF on standard output that has not yet arrived. The next line, which reads standard error, never executes, so the deadlock is never broken. Giving each stream its own reader thread avoids neglecting one stream while reading the other. The PlainJdkRunner example later in this article takes that approach.

If you do not need to tell the two outputs apart, you can merge standard error into standard output with redirectErrorStream(true). That leaves only one pipe to read, so a single thread can handle it.

MergedReadDemo.java
Process process = new ProcessBuilder("sh", "-c", "seq 1 100000 >&2; echo done")
        .redirectErrorStream(true)
        .start();
long lines = process.inputReader().lines().count();
System.out.println("read " + lines + " lines, exit=" + process.waitFor());

It is the same command, but this time it runs to completion. The 100,001 lines are the 100,000 lines written to standard error plus the single done line on standard output.

Output of MergedReadDemo.java
read 100001 lines, exit=0

Redirecting Output the Java Process Will Not Read

If the parent process does not need to refer to or process the child’s output, the simplest option is not to create a pipe that the Java code has to read. The alternatives to the default Redirect.PIPE are as follows.

  • Redirect.INHERIT: the child process writes directly to the parent’s standard output and standard error. (Since JDK 7)

    • The inheritIO() method hands down all three of the parent process’s streams (standard input, standard output, and standard error) at once. (Since JDK 7)

  • Redirect.DISCARD: sends the output to /dev/null. (Since JDK 9)

  • Redirect.to(File): sends the output to a file. (Since JDK 7)

Redirecting the child process’s output
// Send directly to the parent's stdout and stderr
new ProcessBuilder("seq", "1", "100000")
        .redirectOutput(Redirect.INHERIT)
        .redirectError(Redirect.INHERIT)
        .start()
        .waitFor();

// Inherit all three streams at once, including stdin
new ProcessBuilder("seq", "1", "100000")
        .inheritIO()
        .start()
        .waitFor();

// Discard the output to /dev/null
new ProcessBuilder("seq", "1", "100000")
        .redirectOutput(Redirect.DISCARD)
        .redirectError(Redirect.DISCARD)
        .start()
        .waitFor();

// Send to files
new ProcessBuilder("seq", "1", "100000")
        .redirectOutput(Redirect.to(new File("seq-out.log")))
        .redirectError(Redirect.to(new File("seq-err.log")))
        .start()
        .waitFor();

In this environment, all four approaches printed the same 588,895 bytes as the earlier deadlock example and exited normally. Because no child output pipe is created that the parent JVM has to read itself, there is no deadlock from an unread pipe. However, if the output target inherited through INHERIT is itself a pipe, for example if the parent JVM’s standard output is piped to another program, the child’s writes can block or slow down depending on how fast that program reads. When discarding output, specify DISCARD for both standard output and standard error. If you specify only one, the other remains at the default PIPE, and the earlier deadlock can recur. Redirect.to(File) truncates an existing file and overwrites it, so to append on each run use Redirect.appendTo(File).

Closing stdin and EOF

Standard input (stdin), which the Javadoc quoted above warned about alongside the output streams, also defaults to a pipe. As long as the parent JVM keeps the write end of that pipe open, the child process never receives EOF. A command that reads stdin to the end, such as cat run without arguments, keeps waiting even though no more input is coming, and the parent waits in waitFor() for that command to finish. Draining the output pipes does not release this wait.

If you have no input to send, close the write end with process.getOutputStream().close() right after start(), or specify empty input such as redirectInput(new File("/dev/null")). If you do need to send input, close the stream once you have finished writing. zt-exec and Apache Commons Exec, covered later in this article, close the child’s stdin immediately when no input is specified.

The example below runs cat four times, changing only how stdin is handled.

StdinEofDemo.java
Process opened = new ProcessBuilder("cat").start();
System.out.println("stdin left open: finished within 1s=" + opened.waitFor(1, TimeUnit.SECONDS)
        + ", alive=" + opened.isAlive());
opened.destroy();
opened.waitFor();

Process closed = new ProcessBuilder("cat").start();
closed.getOutputStream().close();
System.out.println("stdin closed: finished within 1s=" + closed.waitFor(1, TimeUnit.SECONDS)
        + ", exit=" + closed.exitValue());

Process devNull = new ProcessBuilder("cat")
        .redirectInput(new File("/dev/null"))
        .start();
System.out.println("/dev/null input: finished within 1s=" + devNull.waitFor(1, TimeUnit.SECONDS)
        + ", exit=" + devNull.exitValue());

Process fed = new ProcessBuilder("cat").start();
try (Writer writer = fed.outputWriter()) {
    writer.write("hello\n");
}
String echoed = fed.inputReader().readLine();
System.out.println("closed after writing: echoed=" + echoed + ", exit=" + fed.waitFor());

Here is the output. Only the first run, which left the stream writing to stdin open, failed to finish within one second; the other three received EOF and exited right away.

Output of StdinEofDemo.java
stdin left open: finished within 1s=false, alive=true
stdin closed: finished within 1s=true, exit=0
/dev/null input: finished within 1s=true, exit=0
closed after writing: echoed=hello, exit=0

The first cat was ended with destroy(). Without that cleanup it would stay alive while the remaining examples run. Once the parent JVM exits, however, the write end of the pipe is closed with it, so at that point cat also receives EOF and exits.

The outputWriter() used in the last example was added in JDK 17 together with inputReader(), and it replaces the code that wraps getOutputStream() in an OutputStreamWriter and a BufferedWriter. Called without arguments it uses the charset from the native.encoding system property, and the outputWriter(Charset) method lets you specify one directly.

Using outputWriter()
// Before JDK 17
try (BufferedWriter writer = new BufferedWriter(
        new OutputStreamWriter(process.getOutputStream()))) {
    writer.write("hello\n");
}

// JDK 17 and later
try (BufferedWriter writer = process.outputWriter()) {
    writer.write("hello\n");
}

Timeouts and Termination Policy

Even when you handle all three pipes connected to the child process as described so far, an external command can still hang while waiting on the network or on a lock. That is why a timeout and a termination policy are needed as well. While the command hangs, the thread that called waitFor() blocks until the external command finishes. The waitFor(long, TimeUnit) method, available since JDK 8, merely returns false when the time runs out; it does not terminate the process. At that point you can request graceful termination with destroy() and fall back to destroyForcibly() if the process is still alive after a grace period, or you can kill it right away as the example below does. In the OpenJDK implementation on Linux, destroy() sends SIGTERM and destroyForcibly() sends SIGKILL. A process can catch SIGTERM with a handler and finish up before exiting: deleting temporary files, releasing locks, or cleaning up the child processes it created. SIGKILL terminates the process with no chance to clean up through a handler, so files and descendant processes that were never cleaned up may be left behind. For commands that have something to clean up, it is therefore safer to send SIGTERM first with destroy() and allow a grace period. Note, however, that the process may still be alive briefly after destroyForcibly() returns, so if you need the termination to be complete, confirm it with waitFor(). The example below only runs commands with nothing to clean up, such as sleep, so it kills them right away. It reads the two output streams on JDK 21 virtual threads and puts a one-second limit on waiting for exit.

PlainJdkRunner.java
static void run(String... command) throws IOException, InterruptedException, TimeoutException {
    Process process = new ProcessBuilder(command).start();
    StringBuilder stdout = new StringBuilder();
    StringBuilder stderr = new StringBuilder();
    Thread outPump = Thread.ofVirtual().start(() -> process.inputReader().lines().forEach(l -> stdout.append(l).append('\n')));
    Thread errPump = Thread.ofVirtual().start(() -> process.errorReader().lines().forEach(l -> stderr.append(l).append('\n')));
    if (!process.waitFor(1, TimeUnit.SECONDS)) {
        process.destroyForcibly().waitFor();
        throw new TimeoutException("timed out: " + String.join(" ", command));
    }
    outPump.join();
    errPump.join();
    System.out.println(command[0] + ": exit=" + process.exitValue()
            + ", stdout chars=" + stdout.length() + ", stderr=" + stderr.toString().trim());
}

Here is the result of running four commands, echo hello, seq 1 100000, ls /no-such-dir, and sleep 60, through this method. The 588,895 bytes of output from seq do not cause a hang, the stderr of ls is collected separately, and sleep, which exceeds one second, ends with an exception.

Output of PlainJdkRunner.java
echo: exit=0, stdout chars=6, stderr=
seq: exit=0, stdout chars=588895, stderr=
ls: exit=2, stdout chars=0, stderr=ls: cannot access '/no-such-dir': No such file or directory
timed out: sleep 60

This example targets commands that require no input and create no descendant processes. To use it safely with other commands, you need to take care of the following as well.

  • For commands that read stdin, add the EOF handling from the previous section.

  • Propagate exceptions thrown in the reader threads to the calling thread, and clean up the process and streams on interrupt or timeout as well.

  • The one-second limit in the code above applies only to waitFor(). If you need an overall timeout that also covers start(), the subsequent wait for termination, and join(), implement it separately.

  • If the output is very large, do not collect all of it in a StringBuilder; process it line by line or with a fixed-size buffer.

  • For commands whose output encoding differs from the parent JVM’s native.encoding, specify the encoding with inputReader(Charset).

Cleaning Up Descendant Processes

destroy() and destroyForcibly() send a signal only to the child you created directly. If the command you ran created child processes of its own, those descendants can be left behind. When the child exits first, the descendant processes it created are reparented to PID 1 or to the nearest subreaper process and keep running. A subreaper is a process designated with prctl(PR_SET_CHILD_SUBREAPER); it takes over orphaned descendants in place of PID 1. systemd is the program that, on most Linux distributions, runs first as PID 1 at boot and starts and manages the other services, and it also runs a separate instance called systemd --user for each logged-in user. The systemd --user of a desktop session is one example of a subreaper.

Cases Where Descendants Are Left Behind

There are broadly four situations in which descendant processes are left behind.

  • Commands run through a shell: this is the case when something like sh -c "a | b" or a script such as run.sh is terminated on timeout. The non-interactive shell running the script does not forward the SIGTERM it received to its children, so only the shell exits.

  • Programs run through a wrapper or launcher: with npm run, npm, which runs on Node.js, executes the script via sh -c, so ending the npm process can leave the command the shell started running. LibreOffice’s soffice goes through oosplash to run soffice.bin, which does the actual work, and if you end oosplash with SIGKILL, soffice.bin is not terminated and stays behind. git also runs ssh or git-remote-https as a child during fetch or clone.

  • Programs that turn themselves into daemons: the Gradle daemon, the Kotlin compile daemon, adb start-server, and ssh ControlPersist connections are designed to outlive their caller. Such programs may also leave the original session and group with setsid().

  • Scripts that do not wait for background jobs: if server & is not followed by wait, server keeps running even after the script exits normally. If this descendant keeps the inherited output pipe open, waitFor() returns but the reader of the output never receives EOF.

Let’s reproduce the first case. The code below runs sleep 60 through a shell, calls destroy() after 0.3 seconds, and then prints which of the descendants it looked up before termination are still alive, along with their parent PIDs.

DescendantLeakDemo.java
public static void main(String[] args) throws Exception {
    run("sh", "-c", "sleep 60");
    run("sh", "-c", "sleep 60; echo done");
    run("sh", "-c", "sleep 60 | cat");
    run("bash", "-c", "sleep 60");
}

static void run(String... command) throws Exception {
    Process process = new ProcessBuilder(command).start();
    Thread.sleep(300);
    List<ProcessHandle> descendants = process.descendants().toList();
    process.destroy();
    process.waitFor(1, TimeUnit.SECONDS);
    Thread.sleep(100);

    List<String> left = descendants.stream()
            .filter(ProcessHandle::isAlive)
            .map(p -> name(p) + "(ppid=" + p.parent().map(ProcessHandle::pid).orElse(-1L) + ")")
            .toList();
    System.out.println(String.join(" ", command) + " -> shell alive=" + process.isAlive()
            + ", children=" + descendants.size() + ", left=" + left);
    descendants.forEach(ProcessHandle::destroyForcibly);
}
Output of DescendantLeakDemo.java
sh -c sleep 60 -> shell alive=false, children=1, left=[sleep(ppid=1693)]
sh -c sleep 60; echo done -> shell alive=false, children=1, left=[sleep(ppid=1693)]
sh -c sleep 60 | cat -> shell alive=false, children=2, left=[sleep(ppid=1693), cat(ppid=1693)]
bash -c sleep 60 -> shell alive=false, children=0, left=[]

In all three cases run with sh -c, the shell exited but sleep and cat remained. PID 1693, the parent of the leftover processes, is the systemd --user of this environment. Note that a descendant was left behind even with sh -c "sleep 60", which has only a single command. dash 0.5.12, the /bin/sh on Ubuntu 24.04, ran even the last command received through -c as a child process. bash -c, on the other hand, replaced the shell process with sleep (children=0), so no descendants were left. Because this behavior varies with the kind and version of the shell, it is safer to assume that commands run through a shell leave descendants behind and to have a cleanup method ready.

If the leftover descendant is a command that will finish soon, the problem lasts only as long as that descendant runs. Even so, a task you have already marked as failed on timeout may keep writing files or calling external APIs in the background during that time. The bigger problem is descendants that never finish on their own. The timeout itself means the command did not finish in time, so there is little reason to expect the leftover descendants to finish soon either. Descendants that are not meant to finish, such as servers or daemons, accumulate with each run and consume memory and the process count limit. If they hold on to a port or a file lock, the next run fails with an error like Address already in use.

If a leftover descendant keeps the inherited output pipe open, a read in progress waits for the descendant to close the pipe, which can delay join() as well. The Process implementation in JDK 25 reclaims the output remaining in the pipe and closes the stream when the direct child exits, but whether a wait occurs depends on the ordering between a read already in progress and this exit handling. As in the example above, you can also look up descendants with ProcessHandle.descendants() from JDK 9 and terminate them one by one. But the result is a snapshot, so it cannot reliably clean up processes created between the lookup and the termination, or those reparented after the parent exited. To end the descendants in one go, you have to manage them with OS-level process groups or cgroups.

Terminating by Process Group

For commands whose descendants stay in the same process group, you can run them under setsid to create a new session and group, and then signal the whole group to clean up once the work is done. The JDK API has no facility for creating a process group or signaling one, so this is put together by running the setsid and kill commands.

ProcessGroupKillDemo.java
public static void main(String[] args) throws Exception {
    run("sleep 60 | cat");
    run("(trap '' TERM; sleep 60) & sleep 60");
    run("setsid sleep 60 & sleep 60");
}

static void run(String script) throws Exception {
    Process process = new ProcessBuilder("setsid", "sh", "-c", script).start();
    List<ProcessHandle> descendants = List.of();
    try {
        if (!process.waitFor(1, TimeUnit.SECONDS)) {
            descendants = process.descendants().toList(); // snapshot for checking the result
        }
    } finally {
        killGroup(process.pid()); // the PID of the child, now a group leader via setsid, is the process group ID
        process.waitFor();
    }

    List<String> left = descendants.stream()
            .filter(ProcessHandle::isAlive)
            .map(p -> p.info().commandLine().orElse("?"))
            .toList();
    System.out.println("[" + script + "] pgid=" + process.pid() + ", left=" + left);
    descendants.forEach(ProcessHandle::destroyForcibly);
}

static void killGroup(long pgid) throws Exception {
    signalGroup("TERM", pgid);
    if (!awaitGroupExit(pgid, 5)) {
        signalGroup("KILL", pgid);
        awaitGroupExit(pgid, 5);
    }
}

static boolean awaitGroupExit(long pgid, long seconds) throws Exception {
    long deadline = System.nanoTime() + TimeUnit.SECONDS.toNanos(seconds);
    while (signalGroup("0", pgid)) {
        if (System.nanoTime() > deadline) {
            return false;
        }
        Thread.sleep(100);
    }
    return true;
}

// an exit code of 0 from kill means at least one process in the group received the signal
static boolean signalGroup(String signal, long pgid) throws Exception {
    Process kill = new ProcessBuilder("kill", "-" + signal, "--", "-" + pgid)
            .redirectErrorStream(true)
            .redirectOutput(Redirect.DISCARD)
            .start();
    return kill.waitFor() == 0;
}

The util-linux setsid command only calls fork() when the calling process is already a group leader; otherwise it calls setsid() in its own process and then executes the command. A child created by the JVM is not a group leader, so its PID does not change and process.pid() can be used as the ID of the new group. Passing a negative PID argument to kill sends the signal to the entire process group whose ID is the absolute value. The exception is -1, which means every process you have permission to signal rather than a group. Because kill treats arguments starting with -, such as -9 or -KILL, as signals, -- is placed before the PID so that -295952 is read as a PID. Signal number 0 sends no actual signal and only checks that the target exists and that you have permission, so if kill -0 succeeds, processes remain in the group.

waitFor() only waits for the shell, the direct child, so the shell exiting does not mean the whole group has exited. That is why killGroup() sends SIGTERM, then waits up to 5 seconds while checking with kill -0 whether any processes remain in the group, and sends SIGKILL to the whole group if some are still there. The cleanup is done in finally, always, not only on timeout. Even when the direct child exits normally, descendants started in the background can remain. For example, with a script that only runs sleep 60 &, the shell exits immediately and waitFor() returns true, but the killGroup() in finally cleans up the leftover sleep.

Output of ProcessGroupKillDemo.java
[sleep 60 | cat] pgid=295952, left=[]
[(trap '' TERM; sleep 60) & sleep 60] pgid=295959, left=[]
[setsid sleep 60 & sleep 60] pgid=296018, left=[/usr/bin/sleep 60]

The sleep and cat of the first command were in the same group as the shell, so SIGTERM ended them together. The shell of the second command ended immediately on SIGTERM, but the sleep, which was run to ignore SIGTERM, ended on SIGKILL after the 5-second grace period. Had SIGKILL been skipped on seeing only that the shell had exited, this sleep would have remained. In the third command, the sleep run under setsid moved to a new session and group, so it did not receive the signal sent to the group. When a descendant moves to a new session or a different group with setsid() or setpgid(), this method cannot clean it up. The daemon-style programs seen earlier fall into this category.

Terminating by cgroup

cgroup (control group) is a Linux kernel feature that groups processes into a hierarchy to limit and measure resources such as CPU, memory, I/O, and the number of processes. The memory limits of Docker containers also work through cgroups. A child created with fork() starts out in its parent’s cgroup, so grandchildren and all descendants below them end up in the same cgroup. setsid() changes only the session and the process group, not the cgroup. In cgroup v2, moving a process to another cgroup requires write permission on the cgroup.procs file of the common ancestor of the source and destination cgroups. So a descendant without that permission cannot escape the cgroup. The cgroup the current process belongs to can be seen in /proc/self/cgroup, and the whole tree with the systemd-cgls command.

systemd creates a separate cgroup for each service, and when it stops a service it terminates every process left in that cgroup (KillMode=control-group is the default). With systemd-run --scope, the same approach can be applied to a one-off command. The code below runs the command that escaped the process group earlier as a transient scope unit and stops that unit when the job is done.

SystemdScopeDemo.java
String unit = "job-" + UUID.randomUUID() + ".scope";
Process process = new ProcessBuilder(
        "systemd-run", "--user", "--scope", "--quiet", "--unit=" + unit,
        "sh", "-c", "setsid sleep 60 & sleep 60")
        .redirectError(Redirect.INHERIT)
        .start();
List<ProcessHandle> descendants = List.of();
try {
    if (process.waitFor(1, TimeUnit.SECONDS)) {
        System.out.println("exit=" + process.exitValue()); // a failure of systemd-run also shows up here
    } else {
        System.out.println("cgroup: " + Files.readString(Path.of("/proc/" + process.pid() + "/cgroup")).trim());
        descendants = process.descendants().toList(); // snapshot for checking the result
    }
} finally {
    int stopExit = new ProcessBuilder("systemctl", "--user", "stop", unit)
            .redirectError(Redirect.INHERIT)
            .start()
            .waitFor();
    System.out.println("systemctl stop exit=" + stopExit);
    if (!process.waitFor(10, TimeUnit.SECONDS)) {
        process.destroyForcibly().waitFor();
    }
}

List<String> left = descendants.stream()
        .filter(ProcessHandle::isAlive)
        .map(p -> p.info().commandLine().orElse("?"))
        .toList();
System.out.println("descendants=" + descendants.size() + ", left=" + left);
Output of SystemdScopeDemo.java
cgroup: 0::/user.slice/user-1000.slice/user@1000.service/app.slice/job-0fbacb17-d0c0-47de-a65d-db30d90ecfb1.scope
systemctl stop exit=0
descendants=2, left=[]

When run with --scope, systemd-run waits for the scope unit to start and then replaces its own process with the command, so the PID does not change. The command is still a direct child of the JVM, so you can read its output and wait for it to exit through Process just as in the earlier examples. The cgroup path in the output ends with job-…​scope, which confirms that the command ran in a dedicated cgroup. When systemctl stop was called, both descendants ended, including the sleep that had left the group with setsid. systemctl stop sends SIGTERM first, then SIGKILL to any process that has not exited within the grace period (TimeoutStopSec). The grace period can be set on systemd-run with an option such as -p TimeoutStopSec=5s.

A scope is not tied to a specific process; it stays alive as long as at least one process remains in it. Even if the shell exits first, the scope remains active while a descendant launched in the background is still around, so the unit is always stopped in finally, as in the earlier process group example. If systemd-run fails, for example because it cannot connect to the user bus, it prints an error to stderr and exits with a nonzero exit code, so the example passes stderr through to the parent and prints the exit code. It also checks the exit code of systemctl stop and puts a time limit on waitFor() so that it does not wait forever when cleanup fails. If the command finished first and the scope is already gone, systemctl stop returns exit code 5, meaning the unit does not exist, so real code should treat that case as normal.

This approach has preconditions.

  • Using --user requires the user’s systemd instance to be running. For a service account on a server with no login session, either keep the user instance alive with loginctl enable-linger or use the privileged system instance.

  • Most containers have no systemd inside, so this approach cannot be used there. In that case, consider running a separate container per job and cleaning up at the container level.

  • If you have been delegated permission to manage cgroup v2 directly without systemd, you can create a cgroup for the job, start the command inside it, and then write 1 to the cgroup.kill file to send SIGKILL to every process, including those in child cgroups. This file is supported since Linux 5.14.

Whichever approach you take, the target job must be started inside that cgroup from the beginning, and the permission to move out of it must be managed together with the termination policy.

To sum up, if you control the command you run and it does not become a daemon, the process group approach is enough. If you run arbitrary scripts or commands received from outside, cleaning up by cgroup is the more reliable choice. The libraries in the next section also cut down the code for output handling and timeouts, but they do not solve all of these conditions.

Using zt-exec and Apache Commons Exec

Writing code by hand that covers everything from the previous chapter, the output pipes, EOF on stdin, timeouts, and the termination policy, is not easy. Even with a good understanding of the principles, reimplementing this code in every project is tedious. In practice, using a library that already has this handling built in is the pragmatic choice.

The representative libraries in this area are zt-exec and Apache Commons Exec. Both drain the output pipes on separate threads, and when no input is specified, a class named PumpStreamHandler closes the child process’s stdin so that the child receives EOF. zt-exec’s PumpStreamHandler class is taken from Commons Exec, to the point that the top of its source file says "This file originates from the Apache Commons Exec package". The main differences compared in this article are the shape of the API and the defaults.

zt-exec

zt-exec is a library that ZeroTurnaround, the maker of JRebel, released in 2013 after consolidating the process execution code scattered across its internal projects. Its official name is ZT Process Executor, as the title of the README says, but since the GitHub repository and the Maven artifactId are zt-exec, it is usually called zt-exec. This article uses zt-exec as well. The README explains that the goal is to provide all the functionality of ProcessBuilder and Commons Exec through a single ProcessExecutor class. Its only dependency is slf4j-api, and the minimum Java version is 8.

<dependency>
    <groupId>org.zeroturnaround</groupId>
    <artifactId>zt-exec</artifactId>
    <version>1.13.0</version>
</dependency>

The advantages of zt-exec are as follows.

  • The defaults are safe. Calling execute() with no configuration merges stderr into stdout and discards that output. The deadlock caused by an unread pipe does not happen with the defaults. If you need the output, turn on readOutput(true) and get a string from the result’s outputUTF8(), or pass a function that handles the output line by line to redirectOutput() as a lambda expression.

  • Timeouts are distinguished by exception type. When the wait in execute() exceeds the time set with timeout(), the caller receives a TimeoutException, and the worker thread attempts termination with the stopper. An invalid exit code and a timeout can be told apart by exception type alone. However, there is no guarantee that the stopper has finished terminating the process at the moment the exception is caught. The default stopper only calls destroy(), so on Linux it does not forcibly kill a process that ignores SIGTERM. The termination method can be changed with stopper(). Implement the ProcessStopper interface with a policy that calls destroy(), waits for a grace period, and then calls destroyForcibly().

  • Exit code checking is declarative. By default every exit code is allowed, and when you specify the allowed range with exitValueNormal() or exitValues(0, 1), any other value raises an InvalidExitValueException. The exception object’s getResult() gives access to the output up to that point, which is handy for logging the error message.

  • Auxiliary features are simple to configure. start().getFuture() runs the process asynchronously. However, start() ignores the timeout() setting, so you must limit the wait with something like Future.get(timeout, unit) and handle cancellation and termination separately on timeout. A timeout from get() alone does not terminate the process. destroyOnExit() requests termination of the child process from a JVM shutdown hook. redirectOutputAsInfo() can forward the output to an SLF4J logger, but this method is deprecated and redirectOutput(Slf4jStream.of(logger).asInfo()) is the recommended API.

Here is an example that runs the four commands from the previous chapter with zt-exec.

ZtExecRunner.java
static void captureOutput() throws IOException, InterruptedException, TimeoutException {
    ProcessResult result = new ProcessExecutor()
            .command("echo", "hello")
            .readOutput(true)
            .exitValueNormal()
            .execute();
    System.out.println("output=" + result.outputUTF8().trim() + ", exit=" + result.getExitValue());
}

static void largeOutput() throws IOException, InterruptedException, TimeoutException {
    AtomicLong lines = new AtomicLong();
    ProcessResult result = new ProcessExecutor()
            .command("seq", "1", "100000")
            .redirectOutput(line -> lines.incrementAndGet())
            .timeout(3, TimeUnit.SECONDS)
            .execute();
    System.out.println("seq lines=" + lines.get() + ", exit=" + result.getExitValue());
}

static void timeout() throws IOException, InterruptedException {
    try {
        new ProcessExecutor()
                .command("sleep", "60")
                .timeout(1, TimeUnit.SECONDS)
                .execute();
    } catch (TimeoutException e) {
        System.out.println("timeout: " + e.getMessage());
    }
}

static void exitValue() throws IOException, InterruptedException, TimeoutException {
    try {
        new ProcessExecutor()
                .command("ls", "/no-such-dir")
                .readOutput(true)
                .exitValueNormal()
                .execute();
    } catch (InvalidExitValueException e) {
        System.out.println("exit=" + e.getExitValue() + ", output=" + e.getResult().outputUTF8().trim());
    }
}

Here is the output. The message of the TimeoutException contains the executed command and the time limit, and the error message from ls is read with outputUTF8() thanks to the default that merges stderr into stdout. I also confirmed that no sleep process was left behind after the timeout.

Output of ZtExecRunner.java
output=hello, exit=0
seq lines=100000, exit=0
timeout: Timed out waiting for Process[pid=294823, exitValue="not exited"] to finish, timeout: 1 second, executed command [sleep, 60]
exit=2, output=ls: cannot access '/no-such-dir': No such file or directory

There are caveats as well. With readOutput(true), the thread reading the output pipe stores every byte it reads in a ByteArrayOutputStream. When the result is built, toByteArray() copies it once more, and calling outputUTF8() creates a string as well. Since ByteArrayOutputStream doubles its size when the buffer fills up, the internal buffer alone can grow larger than the output, and a string containing characters outside Latin-1, such as Korean, uses 2 bytes per character. So peak memory usage can exceed twice the size of the output.

If you do not need to keep the entire output, you can leave readOutput off and pass a line-by-line handler to redirectOutput(), as in largeOutput() above. The buffer size in this approach depends on the length of the longest line rather than the total output. Huge output with no line breaks still takes a lot of memory, and if the handler is slow, the child process’s output backs up as well. Such output is better sent to a file or to an OutputStream that works with a fixed-size buffer.

This example used SLF4J API 2.0.17, so with no implementation present, three lines of warnings including "No SLF4J providers were found" appeared during initialization. The dependency declared in the POM of zt-exec 1.13.0 is SLF4J API 1.7.32, and the warning text in that version is different. If you do not need logging, you can add the slf4j-nop that matches the SLF4J API you use to silence the warning.

Apache Commons Exec

Apache Commons Exec is a library that composes command execution, output stream handling, and timeouts from DefaultExecutor, PumpStreamHandler, and ExecuteWatchdog, respectively. Version 1.6.0, used in this article, runs on Java 8 or later and has no external dependencies. DefaultExecutor and ExecuteWatchdog are created with a builder API, and the timeout is specified as a Duration. The old constructors are marked deprecated.

<dependency>
    <groupId>org.apache.commons</groupId>
    <artifactId>commons-exec</artifactId>
    <version>1.6.0</version>
</dependency>

Here is an example that runs the same four commands with 1.6.0. The PumpStreamHandler class in charge of stream handling writes to System.out and System.err by default, so to collect the output you pass your own OutputStream instances as below.

CommonsExecRunner.java
static void run(String... command) throws IOException {
    CommandLine cmdLine = new CommandLine(command[0]);
    cmdLine.addArguments(Arrays.copyOfRange(command, 1, command.length), false);
    ByteArrayOutputStream stdout = new ByteArrayOutputStream();
    ByteArrayOutputStream stderr = new ByteArrayOutputStream();
    ExecuteWatchdog watchdog = ExecuteWatchdog.builder().setTimeout(Duration.ofSeconds(1)).get();
    DefaultExecutor executor = DefaultExecutor.builder().get();
    executor.setStreamHandler(new PumpStreamHandler(stdout, stderr));
    executor.setWatchdog(watchdog);
    try {
        int exitValue = executor.execute(cmdLine);
        System.out.println(command[0] + ": exit=" + exitValue + ", stdout bytes=" + stdout.size());
    } catch (ExecuteException e) {
        System.out.println(command[0] + ": exit=" + e.getExitValue()
                + ", killed by watchdog=" + watchdog.killedProcess()
                + ", stderr=" + stderr.toString().trim());
    }
}
Output of CommonsExecRunner.java
echo: exit=0, stdout bytes=6
seq: exit=0, stdout bytes=588895
sleep: exit=143, killed by watchdog=true, stderr=
ls: exit=2, killed by watchdog=false, stderr=ls: cannot access '/no-such-dir': No such file or directory

On Linux, DefaultExecutor throws an ExecuteException for a nonzero exit code by default. The exit codes that are not treated as errors can be changed with setExitValue() or setExitValues(), and the check can be turned off with setExitValues(null). The sleep in the output above ended from the SIGTERM sent by the watchdog, so its exit code is 143. In this example, after catching the exception, watchdog.killedProcess() is used to tell whether the watchdog intervened. However, this method indicates whether destroy() was called and does not guarantee that the process actually terminated. The ExecuteWatchdog in 1.6.0 only calls destroy() on timeout and has no API for escalating to a forcible kill. If you need a forcible kill, you have to write separate code that finds the processes left after the timeout through ProcessHandle and calls destroyForcibly(). If the command ignores SIGTERM, execute() may keep waiting. Conversely, if a command that handles SIGTERM exits with an allowed code, execute() may return without an exception even after the timeout. So killedProcess() must be checked on the normal return path as well.

Maintenance Status and Recommendation

Below is the status of the two libraries as of 2026-09-06, checked against the version list on Maven Central and each project’s change log.

Item zt-exec Apache Commons Exec

Latest version

1.13.0 (2026-07-10)

1.6.0 (2025-11-25)

Previous releases

1.12 (2020-09-02), 1.11 (2019-07-05)

1.5.0 (2025-05-16), 1.4.0 (2024-01-01), 1.3 (2014-11-02)

Releases in the last 5 years

1

3

Minimum Java version

8

8

Dependencies

slf4j-api

None

Default output handling

Discarded (stderr merged into stdout)

Written to System.out and System.err

Timeout notification

TimeoutException

Checked with watchdog.killedProcess() (ExecuteException if the exit code is a failure)

Nonzero exit code

Allowed by default, InvalidExitValueException when restricted with exitValues()

ExecuteException by default

Termination on timeout

destroy(), replaceable with stopper()

destroy() only

Of the two, I recommend zt-exec. Its API design has three advantages. It distinguishes timeouts from exit code errors by exception type, it takes a single method call to receive the output as a string, and its default settings avoid the deadlock caused by not reading the output pipe. Both libraries use destroy() as the default termination request. zt-exec lets you add a forcible termination policy with stopper(), but Commons Exec requires separate code.

Process Creation on Linux

Java developers running JDK 25 on Linux today can leave the jdk.lang.Process.launchMechanism option unset and keep the default, POSIX_SPAWN. It has been the default since JDK 13. To see why, this chapter first reviews the three ways Linux offers to create a process and the history of the JDK default, then looks in turn at how POSIX_SPAWN works, the memory cost and launch latency of the FORK alternative, and finally the VFORK setting to clean up when upgrading the JVM.

Three Ways to Create a Process and the History of the JDK Default

An application’s code for launching an external process passes through the JVM and ends up as a kernel system call or a C library function call. Which function call creates the child process matters because this choice has led to real production failures in the past. The classic case is a JVM with a large heap failing to allocate memory while trying to run a single external command. The 2015 NAVER D2 article Running external processes in Java (NAVER D2, 2015, in Korean) describes a Tomcat server configured with a large heap on which launching an external process failed with a Cannot allocate memory exception. The OpenJDK issue tracker also records this history. The title of JDK-6868160, which landed in JDK 7, is "(process) Use vfork, not fork, on Linux to avoid swap exhaustion". In other words, changing the call that creates the child process was a measure to avoid swap exhaustion. That you can now simply leave the default alone is the result of these problems being solved one after another.

To run another program on Linux, you pair a call that creates a child process with execve(), the system call that replaces that child with another program. The part where you have a choice is the first half, that is, how to create the child. There are three candidates.

  • fork(): a system call that creates a child by duplicating the parent’s address space.

  • vfork(): a system call that creates a child that shares the parent’s address space instead of duplicating it.

  • posix_spawn(): a C library function that bundles child creation and execve() together.

The Linux kernel handles the fork(), vfork(), and clone() system calls with the same function, kernel_clone(), and flags determine what is shared with the parent. In glibc 2.39 on x86-64, the environment used for this article, the fork() function invokes the clone() system call, while the vfork() function invokes the separate vfork() system call directly. Even in the same glibc 2.39, the AArch64 vfork implementation uses the clone() system call, so the call path differs by architecture. The vfork(2) manual describes vfork() as equivalent to clone() with the flags CLONE_VM | CLONE_VFORK | SIGCHLD. These flags appear in the strace output later in this chapter.

Only posix_spawn() sits at a different layer. It is a function specified by POSIX and implemented by the C library, so it has no corresponding system call; internally it picks one of the two methods above, or a clone() with the same flags as vfork(). The JDK does not implement any of the three methods itself; all three call glibc functions. You can confirm this in the dynamic symbols of libjava.so, which contains the JVM’s process creation code.

$ nm -D $JAVA_HOME/lib/libjava.so | grep -E 'posix_spawn|fork'
                 U fork@GLIBC_2.2.5
00000000000125c0 T Java_java_lang_ProcessImpl_forkAndExec
                 U posix_spawn@GLIBC_2.15
                 U vfork@GLIBC_2.2.5

T marks a symbol this file defines, and U marks a symbol it does not define and looks up in another library at run time. What OpenJDK built is the native method forkAndExec; the functions corresponding to the three launch mechanisms are all linked to glibc symbols carrying @GLIBC_ version tags. So even with the same POSIX_SPAWN setting, the actual behavior depends on the libc and its version. A JDK linked against musl rather than glibc gets musl’s implementation in the same place.

The pros and cons of the three methods are as follows.

Method Advantages Disadvantages

fork()

The child’s address space is separate, so it can do preparation work such as cleaning up file descriptors or changing the working directory without corrupting the parent’s memory

Copy-on-write defers copying the pages, but the page tables are copied, so the larger the heap the higher the creation cost, and the call is subject to the kernel’s overcommit check

vfork()

The page tables are not copied, so the creation cost is nearly independent of the parent’s heap size

Behavior is undefined if the child does anything other than execve() or _exit(), and the parent’s calling thread is blocked in the meantime

posix_spawn()

Preparation work is handled in the way the library defines, and since glibc 2.24 it uses CLONE_VM to avoid copying the page tables as well

The preparation work the child can do is limited to what the library supports, and the internal implementation varies with the libc and its version

That said, fork() does not mean the child may call arbitrary functions before execve() either. The fork(2) manual restricts the functions a fork()`ed child in a multithreaded program can safely call before `execve() to async-signal-safe functions. The other threads are not duplicated into the child, but the lock state those threads were holding can remain.

The vfork(2) manual describes the cost of fork() as "the time and memory required to duplicate the parent’s page tables, and to create a unique task structure for the child". This cost grows with the parent’s heap size. In the measurements later in this chapter, running one command with the FORK mechanism took 16 to 20 ms with a 256MB heap but 230 to 240 ms with an 8GB heap. The constraints on vfork() are stated even more strongly in the vfork(2) manual. Modifying any data other than the pid_t variable that holds the return value, returning from the function that called vfork(), or calling any function other than _exit() and execve() results in undefined behavior. Because the child modifies memory and a stack it shares with the parent, the outcome after the parent resumes cannot be predicted.

OpenJDK went through these three methods in turn. It used fork() at first, but when the creation cost became a problem with large heaps, JDK 7 switched the default to vfork(). The JDK, however, did preparation work between vfork() and execve(), closing file descriptors and changing the working directory, which violated the vfork() constraints we just saw. It was a choice that reduced the launch cost at the expense of safety. POSIX_SPAWN, the default since JDK 13, moves this preparation work into a separate process that does not share memory with the parent, avoiding both the copying cost of fork() and the constraint violations of vfork().

The three methods correspond to the values FORK, VFORK, and POSIX_SPAWN of the implementation-specific system property jdk.lang.Process.launchMechanism. These three values are declared in ProcessImpl.java in JDK 25, and on Linux all three can be specified, though VFORK comes with a warning. The history of the default is shown below. The JDK 27 entry is not a GA release result but what has been merged into the development build as of 2026-09-06.

JDK Change on Linux

6

fork() + execve()

7

vfork() becomes the default (JDK-6868160)

12

POSIX_SPAWN added as an optional mechanism (JDK-8212828). Also backported to 11.0.4

13

POSIX_SPAWN becomes the default (JDK-8213192)

25

Specifying VFORK prints a deprecation warning (JDK-8357180)

27

VFORK removed in the development build. Specifying it prints a warning and falls back to FORK (JDK-8357089, CSR)

Why VFORK went away, and how to clean up any remaining setting, is covered in the last section of this chapter. The next section looks at how POSIX_SPAWN, the default since JDK 13, moves the preparation work into a separate process.

How POSIX_SPAWN Works

posix_spawn() bundles child creation, execve(), and the preparation work in between into a single call. It is similar to CreateProcess() on Windows. The JDK uses this function to first launch a small executable called jspawnhelper; jspawnhelper cleans up file descriptors and the working directory according to the settings it receives over a pipe, then `execve()`s the actual command. This structure is described in the comment at the top of ProcessImpl_md.c in the JDK 25 source.

The reason for the extra hop through the helper is the safety problem from the previous section. Doing preparation work such as closing file descriptors and changing the working directory inside jspawnhelper, which was launched by the first execve() and is therefore a separate process that does not share memory with the parent, does not violate the vfork() constraints. The same comment describes this as "moving the preparation work after the first exec to narrow the vulnerable window".

Viewed with strace, a tool that records the system calls a program makes in order, execve appears twice. Below is the process-creation-related portion of the system calls made when ProcessRunner.java, which runs echo hello, is executed with the default settings. Paths have been shortened.

$ strace -f -e trace=clone,clone3,vfork,execve -o trace.txt java ProcessRunner
$ grep -v CLONE_THREAD trace.txt | grep -E 'clone|vfork|execve'
285583 clone3({flags=CLONE_VM|CLONE_VFORK|CLONE_CLEAR_SIGHAND, exit_signal=SIGCHLD, stack=..., stack_size=0x9000}, 88 <unfinished ...>
285603 execve("$JAVA_HOME/lib/jspawnhelper", ["$JAVA_HOME/lib/jspawnhelper", "25+36-LTS", "10:11:13"], ...) = 0
285603 execve("/usr/bin/echo", ["echo", "hello"], ...) = 0

No vfork() system call appeared in this run. posix_spawn() in glibc 2.39 created the child with the clone3 system call, passing the CLONE_VM and CLONE_VFORK flags. Depending on the environment, the clone() system call may be used instead. The posix_spawn implementation in glibc 2.39 falls back to clone() if the clone3 call fails with ENOSYS or EINVAL. CLONE_VM means the child shares the parent’s address space as is, so the heap mappings and page tables are not copied. CLONE_VFORK blocks the parent’s calling thread until the child execve()`s or exits. Thanks to `CLONE_VM, the creation cost is nearly independent of heap size. The measurements later in this chapter confirm this. Running the same program with -Djdk.lang.Process.launchMechanism=FORK invokes the clone system call without the CLONE_VM flag to duplicate the address space, and execve`s `/usr/bin/echo directly without jspawnhelper.

According to the posix_spawn(3) manual, since glibc 2.24 posix_spawn() calls clone() with the CLONE_VM and CLONE_VFORK flags. It gives the child process a separate stack, blocks signals during creation, and resets the child’s handlers, reducing the risks related to the parent’s stack and signal handling. The comment in the JDK 25 source also describes the glibc 2.24 and later approach as the best choice for these reasons, and notes that musl has used this clone() approach as well. The vfork() function itself has not been removed from glibc.

The fact that jspawnhelper is a separate executable can cause trouble when updating the JDK. If you overwrite the files in the JDK path used by a running JVM, the JVM code loaded in memory and the helper version on disk diverge, and launching external commands can fail. JDK-8325621 strengthened the helper’s version check against the background of such mismatches caused by automatic updates. When updating the JDK, it is better not to overwrite the directory a running JVM uses, but to install into a new directory and switch to the new path when restarting the JVM. Helper launch errors can also have other causes, such as permission or installation problems, so check the error message and the installation state first.

The CSR for JDK-8357090 explains that FORK is kept as an alternative so that unexpected problems with posix_spawn() can be worked around. If jspawnhelper-related errors keep recurring and the cause is hard to pin down right away, you can work around them temporarily by adding -Djdk.lang.Process.launchMechanism=FORK to the JVM startup options. This choice accepts the memory burden under the overcommit policy and the launch latency covered in the next section, and it is not a setting that takes effect immediately in a running JVM.

Memory Cost and Launch Latency of FORK

If you choose FORK, you have to take into account the cost of duplicating the parent JVM’s address space. fork() creates a child process by duplicating the parent’s virtual address space as is; Linux defers the actual copying of pages with copy-on-write, but it does copy the page tables. Depending on the kernel’s memory overcommit policy, the kernel may compute in advance the memory the child might use and refuse to create it. This check applies even if the child is immediately replaced by a small program through execve().

The default value 0 of the Linux kernel setting vm.overcommit_memory, exposed through the /proc/sys/vm/overcommit_memory file, is a mode that "refuses only obvious overcommits". The kernel 5.1 code factored free memory, the page cache, free swap, and so on into this decision. If the size fork() required to duplicate the heap mappings exceeded this computed value, the call could be refused with ENOMEM, the error code meaning out of memory.

The kernel has several rules to keep this decision from becoming too strict, and they have been reinforced as versions progressed. Two examples follow.

  • Even on a server that strictly enforces the commit limit with vm.overcommit_memory=2, the commit check at fork() time does not unconditionally add the size of the parent’s entire address space. It counts only the size of the mappings the kernel includes in the commit total, such as private writable mappings.

  • The mm: fix false-positive OVERCOMMIT_GUESS failures patch, merged in kernel 5.2, simplified the mode 0 decision to "does the requested size exceed the sum of total RAM and swap". This relaxed the condition that refused to duplicate large mappings merely because free memory was low.

Even so, FORK does not always succeed. In mode 2, large anonymous private writable mappings such as the Java heap count toward the cost, so process creation fails if the remaining commit limit is exceeded, and even in mode 0 the actual allocation of kernel memory such as page tables can fail.

Even when creation is not refused, the time cost remains. Although the overcommit decision was relaxed, the fact that fork() copies the page tables has not changed. The copying cost varies with factors such as the heap area actually touched and the page size. POSIX_SPAWN, by contrast, shares the address space through CLONE_VM, so there is no such copying. I measured, for each launch mechanism, the time taken to run the true command 30 times with the heap preallocated. The measurement code is SpawnBench.java; I set -Xms and -Xmx to the same value and pre-touched the heap with the -XX:+AlwaysPreTouch option before measuring. Each run averaged 30 iterations of start().waitFor() after 5 warm-up iterations, so the figures include running the command and waiting for it to exit, not just the pure creation time. Below are the values from two runs at the time of writing.

Heap size POSIX_SPAWN (ms/run) FORK (ms/run)

256MB

1.5, 1.2

15.9, 20.2

2GB

1.6, 1.9

120.8, 140.7

8GB

2.0, 2.1

243.5, 231.2

In this measurement, POSIX_SPAWN stayed in the range of about 1 to 2 ms, while FORK slowed down as the heap grew, differing by more than 100 times at 8GB. This does not mean the cost is exactly proportional to heap size or that this ratio appears in every environment. A re-measurement on 2026-09-06 confirmed the same trend, but the absolute times differed.

In the FORK mechanism, dup_mmap() in Linux 6.17, which duplicates the address space, copies the mappings and page tables while holding the parent process’s address space lock (mmap_lock) in write mode. Meanwhile, other threads attempting memory mapping changes that need this lock have to wait. For a web application that runs external commands synchronously while handling requests, this cost can add to the response time. Unless you are temporarily working around a jspawnhelper problem, there is currently no reason to accept these drawbacks and change the default to FORK.

Removing the VFORK Setting When Upgrading the JVM

When upgrading the JVM to version 25 or later, it is best to remove any remaining -Djdk.lang.Process.launchMechanism=VFORK setting. The places where this setting may still linger are JVM launch scripts, JAVA_OPTS in Dockerfiles, and application server startup options. JDK 25 prints a warning, and with the change merged into the JDK 27 development build the setting is replaced by FORK, which can increase launch latency with large heaps.

Because vfork() was the default on Linux from JDK 7 through 12, some projects wrote this property explicitly to preserve the previous behavior when the default changed to POSIX_SPAWN in JDK 13. Even in environments where glibc was older than 2.24 at the time, there was no need to specify this value. According to the comment in the JDK 13 source, posix_spawn() in glibc 2.4 through 2.23 chooses between fork() and vfork() depending on the call arguments, and the way the JDK calls it meets the conditions for using vfork(). From glibc 2.24 onward, it changed to the clone()-based implementation with the separate stack and signal handling described earlier. There is no reason to specify this value today, so remove the property and return to the default.

JDK 7’s choice of vfork() as the default was controversial even at the time. vfork() is not a standard function. The STANDARDS section of the vfork(2) manual says "None", and the HISTORY section notes that this function, which appeared in 3.0BSD, was marked OBSOLETE in POSIX.1-2001 and had its specification removed in POSIX.1-2008. 4.4BSD made vfork() simply identical to fork(), and Linux also behaved the same as fork() until around 2.2.0-pre6, becoming an independent system call from 2.2.0-pre9. Still, what went away is not the kernel’s vfork() but the JDK’s VFORK mode. It was not because it was dropped from the standard, nor because the kernel changed, but because, as we saw earlier, the work the JDK itself did between vfork() and execve() was unsafe.

Specifying this option in JDK 25 prints the following warning to standard error.

$ java -Djdk.lang.Process.launchMechanism=VFORK ProcessRunner
VFORK MODE DEPRECATED
The VFORK launch mechanism has been deprecated for being dangerous.
It will be removed in a future java version. Either remove the
jdk.lang.Process.launchMechanism property (preferred) or use FORK mode
instead (-Djdk.lang.Process.launchMechanism=FORK).

The examples and measurements in this article are based on JDK 25. When applying them to older JDKs, check the launch mechanism and the range of API support separately. For example, OpenJDK 11u can specify POSIX_SPAWN as an option from 11.0.4 on, but its default is vfork(), and the Linux implementation in OpenJDK 8u does not support this option as of 2026-09-06.

Conclusion

Code that launches external processes must be designed with output handling, timeouts, and a termination policy together. If you do not need to process the output, use Redirect.DISCARD or inheritIO(); if you need stdout and stderr separately, read both streams concurrently. For commands that read stdin, close the stream after writing all the input so the command receives EOF. After a timeout, take care not only of the termination request but also of whether to kill forcibly and of resource cleanup.

zt-exec and Commons Exec reduce this code. Compare their defaults for output and exit codes, how they signal a timeout, and their dependencies, and pick the one that fits your project. Neither library guarantees forcible termination with the default destroy() alone, and for cleaning up descendants you can send a signal to the processes remaining in the same group or terminate by cgroup.

With JDK 25 on Linux, keeping the default POSIX_SPAWN is the safe choice.

References

IEEE 754 Floating-Point Errors and Alternatives in Java

Why storing 0.1 in a Java double causes a rounding error in the 52-bit IEEE 754 fraction, two failures it caused (NEIS grading and the Patriot missile), and how to choose between BigDecimal, integer minor units, and libraries such as Java Money.

Java’s double and float types store values in binary according to the IEEE 754 floating-point standard. In this format, the decimal number 0.1 cannot be represented exactly. In binary, 0.1 is an infinitely repeating fraction, so fitting it into a finite number of bits requires rounding it to the nearest representable value. This article explains why such errors occur and when they become dangerous, using real failures as examples, and then summarizes the alternatives available in Java.

1. Floating-Point Errors in Code

You can see floating-point errors quickly in jshell, the interactive tool included with the JDK since JDK 9.

Adding decimals
jshell> 0.1 + 0.2
$1 ==> 0.30000000000000004

jshell> 0.1 + 0.2 == 0.3
$2 ==> false

jshell> 1.03 - 0.42
$3 ==> 0.6100000000000001

Adding 0.1 ten times does not give 1.0.

Adding 0.1 ten times
jshell> double sum = 0;
sum ==> 0.0

jshell> for (int i = 0; i < 10; i++) { sum += 0.1; }

jshell> sum
sum ==> 0.9999999999999999

The BigDecimal constructor reveals the value actually stored for the double literal 0.1.

The actual value of the double literal 0.1
jshell> new BigDecimal(0.1)
$1 ==> 0.1000000000000000055511151231257827021181583404541015625

I typed 0.1 at the jshell prompt, but the stored value is greater than 0.1 by about 5.55 × 10-18.

This error is not specific to Java. JavaScript’s number type also uses the IEEE 754 binary64 format, so entering the same expressions in the Chrome DevTools console gives the same results.

Floating-point errors shown in the Chrome DevTools console

The next section follows the bit layout of double to show why these values are stored.

2. Bit Layout of double (IEEE 754 binary64) and How 0.1 Is Stored

IEEE 754 is a standard that defines several floating-point formats. Java uses two of them.

Format 1985 name Size Java type

binary32

single

32 bits (sign 1 / exponent 8 / fraction 23)

float

binary64

double

64 bits (sign 1 / exponent 11 / fraction 52)

double

The first version of the standard, IEEE 754-1985, called these two formats single and double. The 2008 revision renamed them binary32 and binary64, and these names carry over to the current standard, IEEE 754-2019. The standard also defines binary16, binary128, and the decimal formats decimal32, decimal64, and decimal128.

The binary formats of IEEE 754 convert a real number to binary scientific notation and then store it as bits. In decimal scientific notation, you shift the decimal point so that the integer part is a single digit from 1 to 9, and express the number of places shifted as a power of 10. For example, 0.001 is written as 1.0 × 10-3. Binary works the same way: you multiply or divide by 2 until the integer part is a single digit, that is, until the value is at least 1 and less than 2. The only nonzero digit in binary is 1, so the normalized result always takes the form ±1.fraction × 2exponent. For example, binary 0.011 (decimal 0.375) is written as 1.1 × 2-2.

A double (binary64) stores this normalized value in 64 bits according to the following rules.

  • The sign (s) is stored in the first bit. 0 means positive and 1 means negative.

  • The exponent (e) is stored in 11 bits as the actual exponent plus 1023. This lets exponents from -1022 to 1023 be stored as unsigned integers from 1 to 2046; reading the value subtracts 1023 again. Of the values 0 to 2047 that 11 bits can hold, the two at either end are reserved: 0 represents zero and numbers very close to zero, and 2047 represents infinity and NaN.

  • The fraction (f) stores only the part after the binary point of 1.fraction, in 52 bits. The integer part of a normalized number is always 1, so that 1 is omitted.

  • If the part after the binary point exceeds 52 bits, it is rounded to the nearest value.

The figure below shows 0.1 stored according to these rules. The 64 bits are divided into a 1-bit sign, an 11-bit exponent, and a 52-bit fraction.

Bit layout of IEEE 754 double and how 0.1 is stored

Let’s follow the conversion in the lower half of the figure. Multiplying 0.1 by 2 repeatedly gives 0.2, 0.4, and 0.8, and on the fourth step it becomes 1.6, which is the first value at least 1 and less than 2. Since multiplying by 2 four times gave 1.6, 0.1 = 1.6 ÷ 24 = 1.6 × 2-4. The sign bit is therefore 0, and the exponent field is -4 + 1023 = 1019. The fraction is the problem. Converting 0.6 to binary gives 0.1001 1001 1001…, with 1001 repeating forever. Fitting it into 52 bits requires rounding. The bits following the first 52 are 1001…, so the discarded part is more than half of the last bit’s weight. The value is therefore rounded up, and the last four bits of the fraction change from 1001 to 1010. As a result, what gets stored is not 0.1 but a number very slightly greater than 0.1. The value 0.1000000000000000055511151231257827021181583404541015625 seen earlier with new BigDecimal(0.1) is exactly this rounded value.

You can see the stored bits in hexadecimal in jshell.

The 64 bits that store 0.1
jshell> Long.toHexString(Double.doubleToLongBits(0.1))
$1 ==> "3fb999999999999a"

One hexadecimal digit is 4 bits. The first three digits, 3fb, are the 12 bits of the 1-bit sign and the 11-bit exponent combined: the sign bit is 0 and the exponent field is 0x3fb, or 1019 in decimal. The remaining thirteen digits are the 52-bit fraction. The digit 9 (1001) repeats twelve times, and only the last digit is rounded up to a (1010).

Not every decimal fraction has this error. Fractions whose denominator is a power of 2, such as 0.5 (1/2), 0.25 (1/4), and 0.75 (3/4), are finite binary fractions, so they are stored exactly as long as they need no more than 53 significant bits. In contrast, a decimal fraction whose reduced denominator contains a factor of 5, such as 0.1 = 1/10, becomes an infinitely repeating binary fraction and cannot avoid rounding error. For more detail on the standard, see the Wikipedia article on IEEE 754.

3. Failures Caused by Binary Representation Errors

In fields that deal with approximations, such as data from scientific experiments, these errors often fall within an acceptable range. But where values must be exactly equal, such as money and grades, or where errors accumulate across calculations, even a tiny error becomes a critical bug. Let’s look at two cases, one in Korea and one abroad, where such errors led to real failures.

3.1. The 2011 NEIS Grading Error

In July 2011, a grade-processing error in NEIS, the next-generation National Education Information System of South Korea, changed the end-of-semester class ranks of about 29,000 students at 823 high schools nationwide. It happened just before the early-admission season for universities, so the impact was large.

NEIS computed a semester total by adding written exam scores and performance assessment scores. The ranks of students with the same total were decided by a tie-breaking rule that each school chose from 45 options (Yeongnam Ilbo article, in Korean). For example, at a school that gives priority to written exam scores, student A with 65 on the written exam and 25 on performance assessments ranks ahead of student B with 60 and 30, even though both have a total of 90. Detecting ties is an important step in computing ranks; schools even had to choose a tie-breaking rule in advance.

NEIS was programmed to display totals to 16 decimal places, and some values had a stray '1' appearing irregularly in those decimal places. This is the same pattern as 1.03 - 0.42 producing 0.6100000000000001 at the beginning of this article. For example, if two students' totals are computed as 90 and 90.00000000000001, the values are not equal, so a direct comparison treats the two students as not tied and sorts the one with the error as having a higher score. As a result, students who should have had the same total were not detected as tied, or the ranks among tied students were reversed (notice from the Ulsan Metropolitan Office of Education, in Korean). The Ministry of Education, Science and Technology explained that the bug was caused by not correcting this calculation error (Kyunghyang Shinmun article, in Korean).

jshell shows that adding the same numbers in a different order can produce a different sum. The following is not NEIS’s actual formula, only an illustration of the principle.

Sums that depend on the order of addition
jshell> 0.1 + 0.2 + 0.3
$1 ==> 0.6000000000000001

jshell> 0.3 + 0.2 + 0.1
$2 ==> 0.6

jshell> (0.1 + 0.2 + 0.3) == (0.3 + 0.2 + 0.1)
$3 ==> false

Both sums add the same three numbers and should be equal, but with double they produce different values and are not detected as a tie.

According to the Ministry’s count, the ranks of 29,007 students at 823 high schools changed, and for 2,416 students at 350 of those schools, the rank grade changed as well (Jeju Ilbo article, in Korean). Schools had to recompute grades with an error-correction program and reissue report cards (Kyunghyang Shinmun article).

The results of an inspection the Ministry’s special task force announced in September of the same year identify the type in which the error occurred. According to the report, while the old NEIS programs were being redeveloped for the next-generation NEIS, the team failed to anticipate arithmetic errors with floating-point (Double) data that arise from the characteristics of the newly installed database (DB2), and this caused the grading errors (Electronic Times article, in Korean). In other words, grades were computed with the DB2 DOUBLE type. The DOUBLE type in Db2 for Linux, UNIX, and Windows is 64 bits with the same range as IEEE 754 binary64, so it has the same kind of error as Java’s double. Some programs were missing the 'error correction code' that was supposed to fix that error, and the ranks were reversed in those programs (Yonhap News article, in Korean).

At the time, network developer Lee Kyung-moon pointed out in a DailySecu article (in Korean) that the problem was using floating point where exact numeric representation was required. He argued that rather than correcting the error, the design should have switched to integer processing or a type that supports arbitrary precision. Integer minor units and BigDecimal, introduced later in this article, are exactly those two options. The incident shows that handling values with decimal arithmetic, such as grades, in floating point can break comparisons such as tie detection.

3.2. The 1991 Gulf War Patriot Missile Failure

The fact that 0.1 cannot be represented as a finite binary fraction has also cost lives. On February 25, 1991, during the Gulf War, an Iraqi Scud missile struck a US Army barracks in Dhahran, Saudi Arabia, killing 28 American soldiers and injuring 99 (US Army article). A Patriot missile battery was deployed at the base to intercept such attacks, but it did not even attempt an interception.

Chapter 1 of the Korean book Software Errors in History (Acorn Publishing, 2014), titled "28 Lives Taken by an Error of 0.000000095," explains the cause of the incident based on the US GAO’s investigation report. The Patriot system counted time in tenths of a second to predict where a target would appear, and converted the count to seconds by multiplying it by 0.1 stored in a 24-bit fixed-point register. Because 0.1 is an infinitely repeating binary fraction, everything beyond 24 bits was truncated, and an error of about 0.000000095 seconds accumulated every tenth of a second. After 20 hours of continuous operation, the computed time was 0.0687 seconds behind the actual time; after about 100 hours, as with the battery at the time of the incident, it was 0.3433 seconds behind. By the GAO report’s calculation, this timing error shifted the range gate where the system looked for the target by 687 meters. The tracking radar searched an area that far from the missile’s actual position. The radar found no target there, so no interceptor was launched, and the Scud hit the barracks.

This system used fixed-point arithmetic, not floating point like double. But the principle is the same as in this article: 0.1 could not fit into a finite number of binary digits, and the error accumulated.

4. Alternatives in Java for Avoiding Decimal Arithmetic Errors

This section introduces two ways to avoid the errors of double arithmetic: using BigDecimal, and handling amounts as integers in the smallest unit.

4.1. Calculating with BigDecimal

The standard way to handle decimal fractions exactly in Java is java.math.BigDecimal. Why You Should Never Use Float and Double for Monetary Calculations also recommends BigDecimal instead of floating-point types for monetary calculations.

Addition and subtraction with BigDecimal
jshell> new BigDecimal("0.1").add(new BigDecimal("0.2"))
$1 ==> 0.3

jshell> new BigDecimal("1.03").subtract(new BigDecimal("0.42"))
$2 ==> 0.61

There is one catch. If you pass a double to the constructor, the rounding error introduced when the decimal fraction was stored in the double carries over into the BigDecimal. As shown earlier, new BigDecimal(0.1) holds the approximation stored in the double, not 0.1. To avoid this error, use the constructor that takes a string, or BigDecimal.valueOf(). BigDecimal.valueOf(0.1) builds the value from the string "0.1" returned by Double.toString(0.1), so it is 0.1. However, valueOf() cannot remove an error that already exists in a value computed as a double. BigDecimal.valueOf(0.1 + 0.2) is 0.30000000000000004.

Using BigDecimal does not decide the rounding policy for you. A division whose decimal expansion does not terminate, such as 1 divided by 3, throws an exception unless you specify the scale and the rounding mode.

Non-terminating division with BigDecimal
jshell> new BigDecimal("1").divide(new BigDecimal("3"))
|  Exception java.lang.ArithmeticException: Non-terminating decimal expansion; no exact representable decimal result.

jshell> new BigDecimal("1").divide(new BigDecimal("3"), 10, RoundingMode.HALF_UP)
$1 ==> 0.3333333333

java.math.RoundingMode, used here, is an enum for choosing how to round: whether to round 0.5 up (HALF_UP), to the nearest even number (HALF_EVEN), or to always discard the fraction (DOWN). For example, rounding 2.5 and 3.5 to integers gives 3 and 4 with HALF_UP, 2 and 4 with HALF_EVEN, and 2 and 3 with DOWN. What BigDecimal eliminates is the error that comes from binary representation. How many decimal places to keep and how to round them is a rule the team must decide and follow according to the needs of the domain.

Joshua Bloch makes the same recommendation in Item 60 of Effective Java, 3rd edition, "Avoid float and double if exact answers are required." The item suggests int and long as well as BigDecimal; the section on choosing between them below covers which to pick.

4.2. Calculating with Integers

Handling amounts as integers in the smallest unit, such as won or cents, is also common. The API of the payment service Stripe is an example. The amount attribute of the Charge object is an integer, described as follows.

A positive integer representing how much to charge in the smallest currency unit (e.g., 100 cents to charge $1.00 or 100 to charge ¥100, a zero-decimal currency).

— Stripe API Reference
The Charge object

This is especially practical for the Korean won, where amounts below 1 won have little value, as long as it does not conflict with business rules. For the same reason, existing systems in Korea often define won amounts as integer columns in the database. So even if the application calculates with BigDecimal, there is at least one conversion to an integer right before the value is stored. An example from the Woowa Brothers tech blog post Writing Test Code with Spock (in Korean) shows this approach. It uses BigDecimal for rounding, but the amount parameter amount and the return type are long.

calculate method from the Spock example
public static long calculate(long amount, float rate, RoundingMode roundingMode) {
    return BigDecimal.valueOf(amount * rate * 0.01)
            .setScale(0, roundingMode).longValue();
}

The amount itself stays an integer, and right after the rate calculation, which produces a fraction, the method returns to an integer through a BigDecimal with an explicit rounding mode.

However, once a floating-point type is mixed in, as in amount * rate * 0.01, the calculation becomes floating-point arithmetic. The long value is converted to float and the multiplication is done in float, losing precision, so errors can occur even though the amount parameter and the return type are integers.

Limits of the calculate method
jshell> calculate(16_777_217L, 100f, RoundingMode.HALF_UP)
$1 ==> 16777216

jshell> calculate(671_089L, 50f, RoundingMode.HALF_UP)
$2 ==> 335544

The first call multiplies by 100%, so it should return 16,777,217 unchanged, but the result is 1 won short. A float has 24 bits of precision (a 23-bit fraction plus the implicit leading bit), so it cannot represent every integer above 224 (16,777,216); 16_777_217L becomes 16,777,216 the moment it is converted to float. The second call should return 335,545, which is 50% of 671,089 won (335,544.5 won) rounded, but it is 1 won short. The amount is less than 224, but the product with the rate, 33,554,450, exceeds 224 and is stored in a float as 33,554,448. Even if the inputs stay integers, a single floating-point step in the middle of the calculation brings the error back.

If you modify the quoted example to move the rate into BigDecimal as well, the result is 16,777,217 as expected.

Version with the multiplication moved into BigDecimal
public static long calculate(long amount, BigDecimal percent, RoundingMode roundingMode) {
    return BigDecimal.valueOf(amount)
            .multiply(percent)
            .movePointLeft(2)
            .setScale(0, roundingMode)
            .longValueExact();
}
Result of the fixed calculate method
jshell> calculate(16_777_217L, new BigDecimal("100"), RoundingMode.HALF_UP)
$1 ==> 16777217

The longValueExact() call on the last line throws an ArithmeticException if the result exceeds the range of long. In that case, longValue() silently returns only the low-order 64 bits, so a wrong amount could be stored.

Even with integer minor units, eliminating errors requires keeping every intermediate value as an integer or a BigDecimal.

5. Choosing Between BigDecimal and Integer Calculation

For a module where calculation precision matters, I recommend at least keeping double and float out of method parameters and return types. The two alternatives are BigDecimal and integer minor units (long), both covered above. long is better for performance and concise code. BigDecimal is better at preventing developer mistakes.

Item 60 of Effective Java, mentioned earlier, frames the choice the same way. Its conclusion is to use BigDecimal if you want the system to keep track of the decimal point and want control over rounding, accepting that it is less convenient and slower than primitive types. It also offers the alternative of using int or long if performance matters and you can keep track of the decimal point yourself, which is the integer approach from the previous section. It adds a rule of thumb based on the number of digits: use int if the values do not exceed nine decimal digits, long if they do not exceed eighteen, and BigDecimal beyond that.

Peter Lawrey of Chronicle Software lists the drawbacks of BigDecimal in If BigDecimal is the answer, it must have been a strange question.

  • The syntax is unnatural. You chain method calls instead of using operators, which makes formulas harder to read.

  • It uses more memory than double. It keeps creating objects, which produces a lot of garbage.

  • It is much slower for most operations.

The post also includes a JMH benchmark measuring the speed difference. The benchmark repeatedly computes the average of two values, rounded to six decimal places, over an array of 1,024 elements.

Table 1. Lawrey’s JMH benchmark results
Benchmark Throughput (ops/s) Error

doubleMidPrice

123,208.083

±2,109.738

bigDecimalMidPrice

23,638.568

±590.094

The double implementation has about 5.2 times the throughput of the BigDecimal one. These numbers were measured in 2014 and may differ on current JVMs. The comparison is against double, not directly against long.

Lawrey’s conclusion is to use BigDecimal if you do not know how to handle rounding with double or if your project standards require it, but, if you have a choice, not to assume that BigDecimal is automatically the right answer. He cites his own experience: when he is brought in to tune the performance of a financial application, he eventually ends up removing BigDecimal. BigDecimal is not the biggest source of latency at first, but as other bottlenecks get fixed, it eventually becomes the slowest part. In the comments, he adds that trading and financial systems existed before BigDecimal, and that many were built in C++ without a decimal type.

Integer minor units have the advantage of using few system resources. A long value requires no object creation, and its arithmetic is plain CPU integer arithmetic. Memory size can be measured with JOL (Java Object Layout). With the default settings of the 64-bit JDK 25 HotSpot VM, a new BigDecimal("1.03") object is 40 bytes. When the value exceeds the range of long, a BigInteger and an int array are added internally, so the 22-digit new BigDecimal("12345678901234567890.12") takes 112 bytes in total. A long value is 8 bytes.

On the other hand, the long approach requires developers to pay more attention to correctness when a calculation involves division or mixes in double or float. BigDecimal also requires a rounding rule, but with long, an error caused by a missing rule is harder to notice. When the rounding mode is missing from a division that does not terminate, BigDecimal throws an exception as shown earlier, but long division silently discards the fraction. For example, 10 / 3 is 3. And because long uses the basic arithmetic operators, the compiler does not stop floating point from creeping in, as with amount * rate * 0.01 in the calculate() example. The whole system must also follow the same rule about what the smallest unit is, a burden that BigDecimal does not have.

The following table summarizes the comparison.

Table 2. Comparison of BigDecimal and integer calculation
Type Representation Strengths Trade-offs

BigDecimal

Arbitrary-precision integer and a scale

  • Represents decimal numbers exactly, with as many digits as needed

  • No need to decide separately how many places to scale to an integer

  • Higher memory use (40 bytes or more per object in the measurement above)

  • Slower arithmetic

  • Arithmetic through chained method calls instead of operators

long

Integer converted to the smallest unit

  • Low memory use with no object creation

  • Fast, using CPU integer arithmetic

  • Errors occur if double or float slips into intermediate calculations

  • The whole system must follow the smallest-unit rule

Once you decide which type to use, write it down as a simple rule that everyone on the team can understand. For example: "Handle amounts as long (in won) in every layer, compute only ratios with BigDecimal, and truncate with RoundingMode.DOWN." The developers responsible for a system keep changing. You cannot assume that everyone understands the principles behind the errors discussed in this article. A clear and simple rule makes it more likely that newly assigned developers follow the same approach. This keeps decimal arithmetic consistent across all modules and reduces the chance that code with critical errors gets in. So I suggest also judging a candidate rule by whether you can write it down concisely and clearly.

After setting a rule, you also need to fix existing code that violates it to keep things consistent from then on. But code that already works and handles money directly is scary to change. That is why an approach, once established, often stays unchanged. Even if only to prepare for changing existing code someday, separate logic that handles sensitive numbers such as money into its own method, like calculate() above. Then you can test it with only inputs and expected results, without a database or external systems. A structure that is easy to test catches not only decimal errors but also logic errors in newly added code early.

6. Domain-Specific Libraries for Money and Units

The two approaches above are choices about how to hold a number. On top of them, there are domain-specific libraries that wrap values with units, such as money, time, and physical quantities, in dedicated types. Money libraries hold the number internally as either a BigDecimal or integer minor units. Java Money is a representative example.

6.1. Java Money (JSR 354)

Java Money is an API for handling monetary amounts together with currencies, standardized as JSR 354. It represents an amount and its currency together with the MonetaryAmount interface. The JCP proposal considered including it in Java SE 9, but it was never added to the JDK. To use it, you add the API (javax.money) and its reference implementation, Moneta, as separate dependencies.

build.gradle
implementation 'org.javamoney:moneta:1.4.5'

The Money implementation calculates with BigDecimal internally, so 1.03 - 0.42, which gave 0.6100000000000001 with double at the beginning of this article, comes out as exactly 0.61.

Subtraction with Money
MonetaryAmount price = Money.of(new BigDecimal("1.03"), "USD");
MonetaryAmount result = price.subtract(Money.of(new BigDecimal("0.42"), "USD"));
System.out.println(result); // USD 0.61

Arithmetic between amounts in different currencies throws javax.money.MonetaryException. The API stops the mistake of mixing amounts in different currencies with an exception at run time.

Adding amounts in different currencies
price.add(Money.of(100, "KRW")); // MonetaryException: Currency mismatch: USD/KRW

Java Money does not replace the two approaches above; it sits on top of them. Two of the main MonetaryAmount implementations in Moneta map directly to the two numeric representations discussed in this article.

  • Money: Stores the amount as a BigDecimal, and so carries the costs of BigDecimal summarized above.

  • FastMoney: Stores the amount in a single long in minor units. The scale is fixed at five decimal places, so the largest representable amount is about 92 trillion (Long.MAX_VALUE / 105). In exchange for this limit it is fast; its Javadoc says it is 10 to 15 times faster than Money.

Because they share an interface, you can swap Money for FastMoney if speed becomes a problem. Even with a domain type, you still need to decide on the numeric representation, and the trade-offs from the previous section reappear as the choice of implementation.

In my view, whether to adopt Java Money depends on whether you handle multiple currencies. With a single currency, plain BigDecimal or long values are enough; in a system that mixes currencies, having the API catch currency mismatches is worth the extra dependency. For more usage examples, see Java Money and the Currency API on Baeldung.

6.2. Other Domain-Specific Libraries

Java Money is not the only library that wraps numeric representations in domain types.

  • Joda-Money: A money library that predates Java Money. The Money class stores the amount as a BigDecimal with the scale fixed to the currency’s default number of decimal places (2 for the dollar, 0 for the yen), while the BigMoney class allows any scale.

  • Units of Measurement API (JSR 385): Represents physical quantities such as length and mass, together with their units, as types. The API package is javax.measure, and the reference implementation is Indriya. Just as a currency mismatch throws an exception for money, it catches calculations with mismatched kinds of quantities, and when the generic types are fixed, it catches them at compile time. For example, code that adds a Quantity<Mass> to a Quantity<Length> does not compile because the generic types differ. Preventing unit errors is a separate matter from calculating values exactly, however. Values are accepted as Number, so a quantity created from a double keeps its floating-point error.

  • java.time.Duration: An example built into the JDK rather than an extra dependency. It stores a length of time in two fields, seconds (long) and nanoseconds (int). It applies the same technique as integer minor units, in that it does not hold fractional seconds in a single double.

Even if none of these libraries fits, a type you build within your own system can play the same role. In that case too, decide whether to hold the number as a BigDecimal or an integer by the criteria described in this article. And make sure float or double cannot slip into intermediate calculations.