A Java service can have plenty of free heap and still stop working. It runs out of database connections, file descriptors, or threads instead. Each of those has a hard limit. A leak eats that limit one request at a time, and the outage arrives days after the deploy that caused it.
A Java resource leak happens when code acquires a limited resource and fails to release it on some path. The resource can be a connection, a stream or file handle, a thread, or an executor. The garbage collector manages memory, not handles, so leaks surface as limit errors rather than heap growth.
- Connections: pool timeouts; an HTTP or JDBC resource not closed on the error path.
- Streams and files:
Too many open files; oftenFiles.lines/Files.listnever closed. - Threads:
OutOfMemoryError: unable to create native thread; threads started per request. - Executors: thread count climbing in steps; a pool created per call and never shut down.
This guide covers the four resource families that cause most leak outages in Java services. For each one it shows how to see the leak coming from a counter, how to find the code path from a dump, and how to write the code so that release is guaranteed. Connection pools already have their own deep dive in HikariCP “Connection Is Not Available, Request Timed Out”. Here they are one case of a broader pattern.
What is a resource leak in Java, and why doesn’t the GC fix it?
A resource leak is a limited resource that code acquires and does not release on at least one path. The resource might be a pooled connection, an operating system file descriptor, a thread, or an executor’s worker threads. Memory leaks are about bytes. Resource leaks are about counts against a limit, and the limits are much smaller: a pool of 10 connections, a descriptor limit of a few thousand, or a few thousand threads.
The garbage collector does not solve this, for three reasons:
- Pooled resources are still reachable. A leaked JDBC connection is referenced by the pool that lent it, so it is never garbage. It stays checked out.
- Running threads are GC roots. A thread that has started and not finished is never collected. Neither is anything it references, including the executor that owns it.
- Cleanup after collection is late. Some JDK file and socket classes register a cleaner that closes the descriptor once the object is collected. Collection happens when the heap needs it, not when descriptors run out. Finalization, the older mechanism, is deprecated for removal (JEP 421).
Every leak comes down to the same shape: acquire, use, and a release that runs on the happy path only.
What does each Java resource leak look like in production?
Each resource fails against a different limit with a different message. Knowing the message tells you which counter to graph.
| Resource | Limit it exhausts | Typical error | Counter to watch |
|---|---|---|---|
| JDBC connection | Pool maximum | HikariCP: Connection is not available, request timed out after 30000ms | hikaricp.connections.active |
| HTTP client response | The HTTP client’s connection pool | Apache HttpClient 4: Timeout waiting for connection from pool; OkHttp logs A connection to … was leaked. Did you forget to close a response body? | Client pool leased/pending metrics |
| Stream, file, socket | Open file descriptor limit (ulimit -n) | IOException/SocketException with Too many open files | process.files.open vs process.files.max |
| Thread | OS thread limits and memory for stacks | OutOfMemoryError: unable to create native thread | jvm.threads.live |
| Executor | Threads, plus heap for queued tasks | The thread error above, or heap growth from an unbounded queue | jvm.threads.live, executor queue size |
| ThreadLocal on a pool thread | Heap | Old-generation growth; data from one request visible in another | Heap floor after full GC |
Metric names are Micrometer’s (Spring Boot Actuator registers the JVM and process ones by default). Older JDKs word the thread error as unable to create new native thread.
The shape of the counter over time matters as much as its value. A leak ratchets with uptime. It does not fall when traffic falls, and a restart resets it. A load problem rises and falls with traffic. If the counter falls sharply at a full GC, something is failing to close streams and the cleaners are closing them late. That is still a bug, and it will lose the race under a burst.
How do you find which resource is leaking?
Work from the counter to the code. The steps below take minutes once the right metrics exist.
- Find the counter that climbs with uptime. Compare open descriptors, live threads, active pool connections, and the old-generation heap floor over several days. Look for the one that rises and does not fall when traffic drops.
- Check the limit it will hit. Read
/proc/<pid>/limits, the pool maximum, and container or user thread limits, and estimate the time until exhaustion. That is your deadline. - Group the leaked resources by kind. Group descriptors by type and name with
lsof, and group thread names from a thread dump with the numbers collapsed. - Find the acquiring code path. File paths, socket peers, thread names, and leak detection traces point at the code that acquires the resource.
- Audit every acquisition of that kind. The dump names one path. Search for every place that acquires the same resource the same way.
- Fix ownership and verify. Use try-with-resources or a shared, lifecycle-managed executor. Then confirm that the counter stays flat for the period that used to show growth.
PID=$(pgrep -f 'java.*orders-service')
grep 'open files' /proc/$PID/limits # soft and hard limit
ls /proc/$PID/fd | wc -l # descriptors open right now
# What kind? REG = regular files, IPv4/IPv6 = TCP sockets, FIFO = pipes
lsof -p $PID | awk 'NR>1 {print $5}' | sort | uniq -c | sort -rn
# Which files or peers? A path or host repeated thousands of times is the leak
lsof -p $PID | awk 'NR>1 {print $9}' | sort | uniq -c | sort -rn | head
jcmd $PID Thread.print > threads.txt
# Collapse numbers so pool-8123-thread-4 and pool-17-thread-1 group together
grep -o '^"[^"]*"' threads.txt | sed -E 's/[0-9]+/N/g' | sort | uniq -c | sort -rn | head
31744 "pool-N-thread-N" <- ~3,968 executors of 8 threads, never shut down
200 "http-nio-N-exec-N"
57 "Timer-N" <- java.util.Timer created and never cancelled
16 "HikariPool-N housekeeper"
Default thread names are evidence in themselves. Executors.defaultThreadFactory() names threads pool-N-thread-M, with N incremented for every new pool, so a pool number in the thousands means pools are being created at runtime. A java.util.Timer starts a thread named Timer-N. Name your own thread factories after their purpose so the dump reads like a map.
Connections and HTTP responses: the leak on the error path
Connection leaks are the best-known resource leak in Java and the best-instrumented. JDBC pools such as HikariCP have leak detection built in. The HikariCP guide covers the JVM specifics, and Connection Leak vs. Pool Exhaustion covers telling a leak from a slow query in any language.
HTTP clients get less attention. An HTTP response holds a pooled connection until its body is consumed or closed. The classic leak is an early return on a non-2xx status that never touches the body:
public Stock fetch(String sku) throws IOException {
Response response = client.newCall(request(sku)).execute();
if (!response.isSuccessful()) {
throw new InventoryException(response.code()); // body never closed: connection leaked
}
return mapper.readValue(response.body().string(), Stock.class); // string() closes, on this path only
}
public Stock fetch(String sku) throws IOException {
try (Response response = client.newCall(request(sku)).execute()) {
if (!response.isSuccessful()) {
throw new InventoryException(response.code()); // closed by try-with-resources
}
return mapper.readValue(response.body().string(), Stock.class);
}
}
This leak hides until the downstream starts failing. While it returns 200s, the body is read and closed. When it starts returning 503s, every error leaks a connection. Your client pool then empties during the one incident where you most need it. The JDK’s own java.net.http.HttpClient has the same rule for streamed bodies. With BodyHandlers.ofInputStream(), the documentation requires you to read to EOF or close the stream to release the exchange’s resources.
Streams and file handles: why does Java say “Too many open files”?
“Too many open files” means the process hit its file descriptor limit. Files, sockets, and pipes all count against the same limit, so a leak of one kind can break the others: a service leaking file handles eventually cannot accept a TCP connection.
The Stream that holds a file open
Files.lines, Files.list, Files.walk, and Files.find return a java.util.stream.Stream backed by an open file or directory. The Files documentation says they must be used in a try-with-resources statement or a similar structure. Consuming the stream with a terminal operation does not close it. Most developers do not expect that, because other streams hold no resources.
long countRows(Path csv) throws IOException {
return Files.lines(csv).skip(1).count(); // file stays open
}
List<Path> pending(Path dir) throws IOException {
return Files.list(dir).filter(p -> p.toString().endsWith(".csv")).toList(); // directory stays open
}
long countRows(Path csv) throws IOException {
try (Stream<String> lines = Files.lines(csv)) {
return lines.skip(1).count();
}
}
List<Path> pending(Path dir) throws IOException {
try (Stream<Path> entries = Files.list(dir)) {
return entries.filter(p -> p.toString().endsWith(".csv")).toList();
}
}
The wrapper whose constructor throws
try-with-resources closes only the resources it has declared. If you wrap one stream in another in a single declaration and the outer constructor throws, the inner stream was never assigned and is never closed. GZIPInputStream reads the gzip header in its constructor and throws a ZipException on a file that is not gzip. A single corrupt upload leaks one descriptor:
// Leaks the FileInputStream if the GZIP header is invalid
try (InputStream in = new GZIPInputStream(new FileInputStream(file))) { ... }
// Each resource is registered before the next one is constructed
try (InputStream raw = Files.newInputStream(file.toPath());
InputStream in = new GZIPInputStream(raw)) { // raw closed even if this throws
...
}
If descriptors leak, a higher nofile limit only moves the outage later. It also hides the growth from anyone watching error rates. Raise the limit when process.files.open is flat but legitimately high, for example on a proxy that holds many sockets. Never raise it to make a climbing line go away.
Threads and ThreadLocals: what outlives the request?
Resources acquired inside a request should be released inside that request. Threads break that rule by design, because a pool thread lives as long as the process. Anything attached to a long-lived thread lives that long too.
Threads started per request
new Thread(...).start() in a request handler looks harmless in a test. Under production concurrency it creates threads without a bound. If the work blocks on a slow dependency, the threads pile up until the JVM cannot create another one. The error is OutOfMemoryError: unable to create native thread, which is about operating system limits and stack memory, not the heap. Raising -Xmx can even make it worse, because it leaves less memory for thread stacks. new Timer() in a request path has the same problem: each Timer owns a background thread, non-daemon by default, until cancel() is called.
ThreadLocal values on pool threads
A ThreadLocal value lives as long as its thread unless it is removed. On a pool thread, that means forever. Code that sets a per-request value (a user context, a large buffer, a formatter) and never removes it retains that object on every pool thread. Worse, the next request on the same thread can read the previous request’s data.
private static final ThreadLocal<TenantContext> CURRENT = new ThreadLocal<>();
public void doFilter(ServletRequest req, ServletResponse res, FilterChain chain)
throws IOException, ServletException {
CURRENT.set(resolveTenant(req));
try {
chain.doFilter(req, res);
} finally {
CURRENT.remove(); // without this, the context stays on the pool thread
}
}
Executors: why is my thread count climbing in steps?
An executor created inside a method and never shut down is a thread leak with a multiplier. A fixed or scheduled thread pool keeps its core threads alive by default, even when idle. Those threads are running, so they are GC roots, and they reference the executor. Nothing can collect it. Each call leaks the whole pool.
public Report build(List<Long> accountIds) {
ExecutorService pool = Executors.newFixedThreadPool(8); // new pool every call
List<Future<Part>> parts = accountIds.stream()
.map(id -> pool.submit(() -> loadPart(id)))
.toList();
return Report.combine(parts); // pool never shut down
}
There are two correct shapes. If the work needs a long-lived pool, create one, give it bounds and a name, and tie its shutdown to the application lifecycle. If the pool is truly scoped to one operation, make the scope explicit. Since Java 19, ExecutorService implements AutoCloseable, and close() waits for submitted tasks to finish.
private final ThreadPoolExecutor pool = new ThreadPoolExecutor(
8, 8, 0, TimeUnit.MILLISECONDS,
new ArrayBlockingQueue<>(500), // bounded queue
Thread.ofPlatform().name("report-", 1).factory(), // Java 21: readable dumps
new ThreadPoolExecutor.CallerRunsPolicy()); // back-pressure when full
@PreDestroy
void stop() throws InterruptedException {
pool.shutdown();
if (!pool.awaitTermination(30, TimeUnit.SECONDS)) pool.shutdownNow();
}
try (ExecutorService scope = Executors.newVirtualThreadPerTaskExecutor()) {
List<Future<Part>> parts = accountIds.stream()
.map(id -> scope.submit(() -> loadPart(id)))
.toList();
return Report.combine(parts);
} // close() waits for the tasks, then the executor is done
The Executors factory methods also hide their bounds, and each has a different failure mode under load:
| Factory | Threads | Queue | Failure mode under a slow dependency |
|---|---|---|---|
newFixedThreadPool(n) | Fixed at n | Unbounded | Queue grows on the heap; latency grows without errors |
newCachedThreadPool() | Unbounded | None (hand-off) | A new thread for every waiting task; thread exhaustion |
newScheduledThreadPool(n) | n core | Unbounded delay queue | Tasks scheduled per request accumulate if never cancelled |
newVirtualThreadPerTaskExecutor() | One virtual thread per task | None | Cheap threads, but downstream pools (connections) now take the pressure |
A resource without an owner is a resource that nobody releases. Most leaks are an ownership bug, not a typo.
How do you prevent Java resource leaks before production?
Leaks survive testing because tests exercise the happy path, run for seconds, and restart the JVM between suites. Prevention has to target the paths and durations that tests skip.
- Every
AutoCloseableacquired in a method is closed with try-with-resources - Methods that return an open resource say so in the type and name
- No
new Thread,new Timer, orExecutors.new*in request paths - Every
ThreadLocal.sethas aremove()infinally
- SpotBugs
OBL_UNSATISFIED_OBLIGATIONandOS_OPEN_STREAM - Error Prone
MustBeClosedCheckerwith@MustBeClosedon factories - Tests that assert open file and thread counts return to baseline
- Alert on
process.files.open / process.files.maxtrend, not just the value - Alert on
jvm.threads.livegrowth with uptime - Pool leak detection enabled with a measured threshold
Static analysis catches the local cases well: a stream opened and not closed in the same method. It struggles with ownership that crosses methods, such as a resource returned to a caller, stored in a field, or handed to another thread. Those need a review question: who closes this, and on which paths? That question is part of reviewing a pull request for production reliability risks.
Fixes, matched to the resource
| Leak | Fix | What not to do |
|---|---|---|
| JDBC or HTTP connection | try-with-resources around connection, statement, result set, and response; close on error paths | Raise the pool size |
| File or directory stream | try-with-resources around Files.lines/list/walk; declare wrapped streams separately | Raise ulimit -n while the count climbs |
| Threads per request | Submit to a shared, bounded executor | Lower -Xss to fit more threads |
| Executor per call | One application-owned executor, or try-with-resources on Java 19+ | Rely on garbage collection to shut it down |
| ThreadLocal on pool threads | remove() in finally, set at a single entry point | Use larger heaps to absorb it |
How Tomosu helps
A thread dump or an lsof listing shows the leak that is happening now. Tomosu looks at the code for the ones that have not happened yet, across the repository and in each pull request:
- Acquire and release paths: where connections, HTTP responses, file streams, and
Files.*streams are opened, and whether each path, including error and early-return paths, releases them. - Lifetime mismatches: threads, timers, and executors created inside request paths, executors without a shutdown, and
ThreadLocal.setwithout a matchingremove(). - Blast radius: which endpoints, jobs, and consumers reach the leaking path, and which shared limits (pool, descriptors, threads) they compete for.
- Evidence to confirm: which counter to watch and what it should look like after the fix.
These signals roll into the Production Reliability Index, so a change that adds Executors.newFixedThreadPool to a request handler can be caught in review.
Scan your repository with Tomosu →
Key takeaways
- Java resource leaks exhaust counts, not bytes: pool slots, file descriptors, threads. The heap can look healthy right up to the outage.
- The garbage collector does not rescue you. Pooled connections stay referenced, running threads are GC roots, and cleaners run late.
- Most leaks are a release that runs only on the happy path. try-with-resources covers exceptions and early returns.
Files.lines,Files.list, andFiles.walkhold a file or directory open until the stream is closed.- A thread count that climbs in steps of the pool size means an executor is created per call and never shut down.
- Find the counter that ratchets with uptime, group the leaked resources, then audit every acquisition of the same kind.
- Raising
ulimit, pool sizes, or heap only moves the outage later.
Frequently asked questions
What is a resource leak in Java?
A resource leak is a limited resource that code acquires and does not release on some path: a database connection, a file or socket descriptor, a thread, or an executor. The garbage collector manages memory, not these handles, so a leaked resource stays taken until the process runs out and fails.
Doesn't the garbage collector close leaked streams and connections?
Not reliably. Some JDK classes register a cleaner that closes the descriptor after the object becomes unreachable and is collected, but that can happen much later or not before the limit is reached. Pooled connections are still referenced by the pool, and running threads are never collected, so neither is cleaned up by garbage collection.
What causes java.io.IOException: Too many open files?
The process reached its open file descriptor limit. Files, sockets, and pipes all count. The usual cause is a stream, file, or HTTP response that is not closed on some path, often a Stream returned by Files.lines, Files.list, or Files.walk. Raising the ulimit only delays the failure if descriptors keep leaking.
What does OutOfMemoryError: unable to create native thread mean?
The JVM asked the operating system for a new thread and was refused, because of a process or user thread limit or because there was no memory for another thread stack. It is rarely a heap problem. The common cause is code that creates threads or executors per request and never ends them.
Do I need to shut down an ExecutorService?
Yes, if you created it. Pool threads of a fixed or scheduled executor do not time out by default and are not garbage collected while they run, so an executor that is never shut down keeps its threads for the life of the process. Share long-lived executors and shut them down with the application, or use try-with-resources on Java 19 and later.
Can ThreadLocal cause a memory leak in a thread pool?
Yes. A value set in a ThreadLocal lives as long as the thread, and pool threads live as long as the pool. If code sets a value per request and never calls remove(), large objects stay reachable and the next request on that thread can see the previous request's data. Remove it in a finally block.
How can I catch Java resource leaks before production?
Enforce try-with-resources for every AutoCloseable, run static analysis such as SpotBugs or Error Prone that flags unclosed resources, forbid creating executors and threads inside request paths, and add tests that assert open file and thread counts return to their baseline after a workload.
Every resource leak is a question nobody answered in review: who releases this, and on which paths? Tomosu asks it for every acquisition in the codebase. Assess your repository →