The dashboard shows a Node.js service whose memory line only goes up. It restarts every few days, or the platform kills it, and the graph starts climbing again. Before anyone reaches for a bigger heap or a nightly restart, you need to answer two questions: is this actually a leak, and if so, which object in which code path is holding on to the memory?
When Node.js memory keeps growing in production, first check whether the heap floor after each full garbage collection is rising. A rising line on its own is not proof, because V8 lets garbage pile up between collections. If the floor rises under steady traffic, find the layer that grows, then compare heap snapshots to see what is retained.
- Growing
heapUsedfloor: JavaScript objects retained by a Map, listener, timer, or closure. - Growing
arrayBuffers/external: Buffers held by streams, slices, or ignored backpressure. - Only
rssgrows: native memory or allocator fragmentation; heap snapshots will not show it. - Three snapshots under load, then follow the retainer chain to the code that owns the memory.
This guide is a step-by-step investigation for a long-running Node.js service: how to read the growth curve, which process.memoryUsage() field matters, how to take and compare heap snapshots without causing a second incident, and the code patterns that show up at the end of most retainer chains. If your service runs in Kubernetes and the symptom is an OOMKilled pod, start with Kubernetes OOMKilled: Memory Leak, Memory Limit, or Something Else?, which covers heap caps, container limits, and what the cgroup counts. This post goes deeper into the Node.js side.
Is growing Node.js memory actually a leak?
Not necessarily. A memory leak in a garbage-collected runtime is memory that stays reachable after the program no longer needs it. The garbage collector cannot free it because something still points to it. That is different from memory that is simply not collected yet.
V8, the JavaScript engine inside Node.js, is lazy on purpose. Short-lived objects are collected cheaply and often in the young generation. Objects that survive move to the old generation, which is collected by a much more expensive full mark-compact pass. V8 runs that pass when it decides the heap has grown enough, not when your dashboard would like it to. Between full collections, heapUsed climbs with garbage that is already unreachable. heapTotal, the memory V8 has reserved for the heap, can also grow toward the heap limit and stay there.
So the raw memory line can go up for hours in a perfectly healthy process. The signal that matters is the floor: how much heap is still in use right after a full collection.
The easiest way to see the floor is V8’s own GC trace. Start the process with --trace-gc (a V8 flag Node passes through) and read the full collections. It is verbose, so enable it on one instance for a limited period:
# node --trace-gc server.js 2>&1 | grep -E 'Mark-(Compact|sweep)'
[1:0x5a1c] 3600512 ms: Mark-Compact 212.4 (231.0) -> 148.2 (229.5) MB, ...
[1:0x5a1c] 7201877 ms: Mark-Compact 239.8 (258.3) -> 171.9 (256.0) MB, ...
[1:0x5a1c] 10803004 ms: Mark-Compact 266.1 (284.7) -> 195.6 (282.1) MB, ...
# The number after "->" is heap in use after the collection. +24 MB per hour: a leak candidate.
Older Node.js versions label the full collection Mark-sweep instead of Mark-Compact. The number before the arrow is heap in use before the collection, the number after is what survived, and the values in parentheses are the reserved heap. If you already export metrics, the same trend is visible by plotting the minimum of heapUsed over a sliding window of several minutes.
Which memory number is growing: heap, external, or RSS?
A Node.js process has more memory than the V8 heap. process.memoryUsage() reports the layers separately, and knowing which one grows decides every next step. Heap snapshots, the main tool for JavaScript leaks, see only one of them.
| Field | What it measures | If it is the one growing |
|---|---|---|
heapUsed | Live JavaScript objects plus garbage not yet collected | Watch its post-GC floor. A rising floor means retained JavaScript objects. Heap snapshots will find them. |
heapTotal | Heap memory V8 has reserved | Growth toward the limit alone is normal. Read it together with heapUsed. |
external | Memory of C++ objects bound to JavaScript objects | Usually Buffers (see next row). Otherwise native addons that allocate through V8. |
arrayBuffers | Backing stores of ArrayBuffer and Node.js Buffer (also counted in external) | Retained Buffers: stream buffers, sliced Buffers, request bodies held in memory. |
rss | Everything resident for the process: heap, code, stacks, native allocations | Rising while the others are flat means native memory or allocator fragmentation. Snapshots will not help. |
A short periodic log line with these fields, emitted every 15 to 60 seconds, is enough to answer all three questions. If you already use prom-client, its default metrics export heap and resident memory. Check that arrayBuffers or external is included, because that is the layer teams most often forget.
Step by step: how do you investigate a Node.js memory leak in production?
The procedure below works whether the service runs on a VM, in a container, or on a platform that restarts it for you. Steps 1 to 3 use data you probably already have. Steps 4 to 6 need a heap snapshot tool and one instance you can set aside.
- Confirm the shape of the growth. Plot the heap floor after full GCs over at least several hours, next to request rate. Flat floor: stop here. Rising then flat: check that caches have bounds. Rising steadily: continue.
- Find which layer is growing. Use the decision tree above. Only continue with heap snapshots if the
heapUsedfloor is the one rising. - Correlate growth with a driver. Divide growth by request count to get bytes retained per request, and check whether growth follows requests, wall-clock time, unique keys, or a specific deploy (see the table below).
- Capture three heap snapshots. From one instance: after warm-up, after a period of normal load, and after a second period of load.
- Compare the snapshots. In the third snapshot, view the objects allocated between the first and second. They survived a whole load cycle and several full GCs.
- Follow the retainer chain to code. Pick the largest surviving group and read its retainers until you reach a variable that your code owns.
- Fix and verify under the same load. The fix is proven only when the post-GC floor stays flat under the traffic pattern that produced the growth.
Step 3 in practice: what does the growth follow?
This step is cheap and often skipped. It narrows the search before you open a single snapshot. Suppose the floor grows 24 MB per hour while the instance serves 40 requests per second, which is 144,000 requests per hour. That is roughly 170 bytes per request. That is too small for a leaked request body, but it fits a map entry keyed by request ID, or one listener closure per request.
| Growth follows | Suspect first |
|---|---|
| Request count, including at night | Something stored per request: a pending-request map, a per-request listener, a per-request closure in a long-lived array. |
| Wall-clock time, even with no traffic | A timer: setInterval that appends to a structure, or intervals created and never cleared. |
| Number of distinct users, tenants, or URLs | A keyed cache or metric without a bound: a memoization map, a per-tenant client, a metric label with user IDs. |
| Specific endpoint or payload size | Buffers or parsed payloads held by one handler; stream backpressure ignored on large downloads. |
| A particular deploy | Diff the release. The commit investigation workflow applies to memory regressions too. |
Steps 4 and 5: the three-snapshot comparison
A single heap snapshot tells you what is big. It does not tell you what is growing, and big caches that are working as designed dominate it. Comparing snapshots removes that noise. The three-snapshot technique goes one step further: it isolates objects that were created during one load period and were still alive a full load period later.
To take the snapshots from a running process, start it with --heapsnapshot-signal=SIGUSR2 and send that signal, or call v8.writeHeapSnapshot() from a protected admin route. Load the three .heapsnapshot files in the Chrome DevTools Memory panel. Then:
- Select snapshot 3, and in the Summary view change the filter from All objects to Objects allocated between Snapshot 1 and Snapshot 2.
- Sort by Retained Size. Retained size is the memory that would be freed if this object were collected; shallow size is only the object itself.
- Look for constructors whose count grows in proportion to load: your own classes,
Objectshapes with request fields,(closure),Promise,Timeout, or aMap/Arraywith a very large retained size. - Cross-check with the Comparison view between snapshots 2 and 3, sorted by # Delta, to confirm the same types keep growing.
The DevTools heap snapshot documentation describes each view, and the Node.js guide Using Heap Snapshot covers the capture options.
How do you read a retainer chain?
A retainer chain is the path of references from a GC root to the object you selected. The Retainers pane at the bottom of the Memory panel shows it from the object upward. Everything in that chain is a reason the object cannot be collected. Your job is to find the first link that belongs to your code and should not be there.
Three reading rules save most of the time:
- Skip internal links such as
(system),(internal array), and hidden class maps. Keep going until a name you recognise appears: a variable, a property, or a module. - Prefer the shortest distance. An object can have several retainers; the one with the smallest distance from the root is usually the real owner.
- Closures show up as
context. A(closure)retains every variable in the scope it captured, which is how a small callback can keep a large request body alive.
Which code patterns cause most Node.js memory leaks?
Almost every retainer chain ends in one of a handful of patterns. All of them share one property: a structure that lives for the life of the process gains an entry per request, per key, or per tick, and the code that removes the entry runs on only some paths.
| Pattern | What the snapshot shows | Fix |
|---|---|---|
| Pending-request map | A Map of resolvers or callbacks whose count keeps rising | Delete the entry in a finally, on timeout, and on error, not only on success |
| Unbounded cache or memoization | A module-level Map or object keyed by user, URL, or query | A bounded cache (lru-cache with max and ttl), or no cache |
| Listener added per request | Growing listener arrays on a long-lived emitter; MaxListenersExceededWarning | Remove the listener when the request ends, or use { once: true } |
| Timers never cleared | Growing Timeout objects and the closures they capture | Keep the handle; clearInterval/clearTimeout on every exit path |
| Metric label cardinality | Metrics registry objects growing with user IDs or raw paths | Label by route template and status, never by ID |
| Retained Buffers | Growing arrayBuffers; small slices keeping large parent Buffers alive | Buffer.from(slice) to copy what you keep; respect backpressure with pipeline() |
| In-memory session store | Session objects accumulating per visitor | An external session store; the default express-session MemoryStore warns it is not for production |
The pending-request map
This is the pattern in the retainer chain above, and it is common in hand-written RPC clients, WebSocket request/response layers, and queue consumers that wait for replies. The success path cleans up. The timeout and error paths do not.
const pending = new Map(); // lives as long as the process
function call(msg) {
return new Promise((resolve, reject) => {
pending.set(msg.id, { resolve, reject });
setTimeout(() => reject(new Error('timeout')), 5000); // entry is never deleted
socket.send(JSON.stringify(msg));
});
}
socket.on('message', (raw) => {
const reply = JSON.parse(raw);
const entry = pending.get(reply.id);
if (entry) { pending.delete(reply.id); entry.resolve(reply); } // only the success path cleans up
});
function call(msg, timeoutMs = 5000) {
return new Promise((resolve, reject) => {
const timer = setTimeout(() => {
pending.delete(msg.id); // timeout path
reject(new Error('timeout'));
}, timeoutMs);
pending.set(msg.id, {
resolve: (v) => { clearTimeout(timer); pending.delete(msg.id); resolve(v); },
reject: (e) => { clearTimeout(timer); pending.delete(msg.id); reject(e); },
});
try { socket.send(JSON.stringify(msg)); }
catch (err) { pending.get(msg.id)?.reject(err); } // send failure path
});
}
Note what is not the leak here. A promise that never settles is not retained by itself; if nothing references it, it is collected with its callbacks. The leak is the long-lived Map that holds the resolver, and through its closure the caller’s context.
The listener added per request
Node warns when an emitter gets more than 10 listeners for one event, which is the default maxListeners. The warning names the event and the emitter type:
(node:1) MaxListenersExceededWarning: Possible EventEmitter memory leak detected.
11 change listeners added to [EventEmitter]. Use emitter.setMaxListeners() to increase limit
(node:1) MaxListenersExceededWarning: Possible EventTarget memory leak detected.
11 abort listeners added to [AbortSignal]. Use events.setMaxListeners() to increase limit
The fix is almost never setMaxListeners(), which only silences the warning. Run with --trace-warnings to get the stack of the line that added the eleventh listener, then remove the listener when the request is done:
// Broken: one listener per request on an emitter that lives forever
configBus.on('change', () => res.locals.stale = true);
// Fixed: register once at startup, or remove it when the request finishes
const onChange = () => { res.locals.stale = true; };
configBus.on('change', onChange);
res.on('close', () => configBus.off('change', onChange)); // released per request
The unbounded cache
An in-process cache is a leak with good intentions. It looks like the warm-up curve in the first diagram until the key space turns out to be much larger than expected: every user, every URL with a query string, every tenant. Give every cache a size bound and an expiry, and decide what happens when it is full.
// Broken: grows with every distinct (sku, region, currency) ever requested
const cache = {};
// Fixed: lru-cache v10+
import { LRUCache } from 'lru-cache';
const cache = new LRUCache({ max: 10_000, ttl: 5 * 60 * 1000 });
A Prometheus client keeps one time series in memory for every distinct combination of label values. A histogram labelled with userId or the raw request path (/orders/8812) grows with every new value. The retainer chain ends in the metrics registry, not in your business code. Label by route template (/orders/:id) instead.
When is growing memory not a leak?
Some of the most expensive memory investigations end with “nothing is leaking.” Rule these out before you blame the code:
- The heap is growing toward its limit. On a large machine or container, V8 may be allowed a large old generation and will use it before a full collection. The post-GC floor is what counts.
- Traffic or payloads changed. A new client sending 5 MB JSON bodies raises every layer. Growth that tracks payload size and falls with it is load, not a leak.
- Allocator behaviour. Native allocators may keep freed memory mapped for reuse, so
rsscan plateau high after a spike even when the heap has shrunk. With glibc, many threads can mean many arenas; tuning such asMALLOC_ARENA_MAXor a different allocator is sometimes used, but measure before and after. - Worker threads. Each
worker_threadsWorker has its own V8 heap. A pool that grows its worker count grows memory in steps.
A scheduled restart or an aggressive memory-based restart policy keeps the service up, but it also resets the curve before the leak becomes obvious. If you use one as a stopgap, keep one instance running long enough to capture the three snapshots.
How do you capture evidence safely in production?
Heap snapshots are the best evidence and the most disruptive to collect. According to the Node.js v8 documentation, writing a snapshot is synchronous, so it blocks the event loop. It also needs memory about twice the size of the heap at that moment. On a process that is already close to its limit, taking the snapshot can cause the crash you are investigating.
| Method | Cost | Use it when |
|---|---|---|
process.memoryUsage() metrics | Negligible | Always. It answers the “which layer” question. |
--trace-gc | Log volume | For a limited time on one instance to read the post-GC floor. |
--heap-prof (sampling heap profiler) | Low; written when the process exits | To see which call sites allocate the most over a run, in a test or canary. |
Heap snapshot (--heapsnapshot-signal, v8.writeHeapSnapshot()) | Blocks the event loop; ~2× heap memory | On one instance removed from the load balancer, early in the growth curve. |
--heapsnapshot-near-heap-limit=N | Same, at the worst moment | As a last resort to catch the state just before a heap out-of-memory crash. |
Inspector attached via --inspect | Interactive, pauses on snapshot | Staging, or production via an SSH tunnel. Never bind the inspector to a public interface. |
- Take the instance out of rotation first, then send the signal. The load balancer should not route to a process that will freeze for seconds.
- Write to disk-backed storage with enough free space for several snapshots of the full heap size.
- Treat snapshot files as sensitive data. They contain every string in memory: tokens, session IDs, personal data from request bodies. Store and delete them accordingly.
- Capture early. A snapshot at 40% of the heap limit is far safer than one at 90%, and the comparison works just as well.
Most Node.js memory leaks are cleanup code that runs on the success path and nowhere else.
Fixes, matched to the cause
| Cause | Fix | What not to do |
|---|---|---|
| Per-request entries in a long-lived structure | Remove the entry on every exit path: success, error, timeout, client disconnect. Use finally. | Add a periodic sweep that hides a missing cleanup path. |
| Unbounded cache | Bound by count and age; measure hit rate; consider an external cache for large key spaces. | Raise --max-old-space-size so the cache has more room. |
| Listener or timer leak | Register once, or pair every on/setInterval with an off/clearInterval tied to the owner’s lifetime. | setMaxListeners(0) to silence the warning. |
| Retained Buffers | Stream instead of buffering; use stream.pipeline(); copy small slices you keep. | Read whole files or bodies into memory “because they are small.” |
| Native memory or fragmentation | Isolate the addon or allocation pattern; test allocator settings with measurements. | Take more heap snapshots. They cannot see this memory. |
| Heap legitimately too small | Raise the heap limit with headroom, below any container limit. | Treat this as the default fix before the floor is measured. |
Verify with the same measurement that raised the alarm: run the traffic pattern that caused growth and watch the post-GC floor. A fix that makes the raw line look better for an hour proves nothing. If the leak arrived with a recent release, the same investigation pairs well with correlating logs, traces, and a code change. The same “acquire without release” shape appears for other resources in Resource Leaks in Java Services.
How Tomosu helps
A heap snapshot finds the structure that leaked today, in one process. It does not list the other structures written the same way that have not leaked yet. Tomosu analyzes the repository and each pull request for that list:
- Long-lived structures that grow: module-level
Maps, arrays, and caches that gain entries on a request path, and whether anything bounds or removes them. - Cleanup that runs on only some paths: entries deleted on success but not on timeout or error, listeners and timers registered per request without a matching removal.
- Blast radius: which endpoints and consumers reach the growing structure, so a finding on a hot request path is weighed above one in an admin script.
- The evidence to ask for: which memory layer to watch and what the post-GC floor should look like after the fix.
These findings roll up into the Production Reliability Index, so a change that adds a per-request entry to a process-wide map can be discussed in review rather than discovered from a memory graph.
Scan your repository with Tomosu →
Key takeaways
- Rising Node.js memory is not proof of a leak. Watch the heap floor after full garbage collections, not the peaks.
- Decide which layer grows first:
heapUsed,arrayBuffers/external, or onlyrss. Heap snapshots see only the heap. - Correlate growth with requests, time, distinct keys, or a deploy, and estimate bytes retained per request.
- Use three snapshots and view objects allocated between the first and second in the third.
- Follow retainers to the first long-lived container your code owns. That container, not the biggest object, is the fix point.
- Most leaks are cleanup that runs only on the success path: maps, listeners, timers, and caches without bounds.
- Take snapshots from an instance out of rotation, early, and handle the files as sensitive data.
Frequently asked questions
Is it normal for Node.js memory to keep growing?
Some growth is normal. V8 collects garbage lazily, so heapUsed rises between collections and heapTotal can grow toward the heap limit before a full collection runs. What is not normal is a heap floor, measured right after full garbage collections, that keeps rising for hours under steady traffic. That pattern points to a leak.
How do I find a memory leak in a Node.js application in production?
Confirm the post-GC heap floor is rising, identify whether heapUsed, arrayBuffers, or only rss is growing, then take three heap snapshots from one instance at intervals under load. In the third snapshot, look at objects allocated between the first and second, and follow their retainers back to the Map, listener, timer, or closure in your code that holds them.
Why does RSS keep growing when the Node.js heap is flat?
RSS counts all memory the process has resident, not just the V8 heap. If heapUsed and external are flat while RSS rises, suspect native memory from addons, memory held by the allocator after fragmentation, or thread stacks. Heap snapshots will not show this memory, so compare process.memoryUsage() fields before spending time on snapshots.
Is MaxListenersExceededWarning a memory leak?
It is a strong hint, not proof. Node emits it when more than 10 listeners for one event are added to one emitter by default. When the count keeps climbing, code is usually adding a listener per request to a long-lived emitter and never removing it. Run with --trace-warnings to get the stack trace of the code that added the listener.
Is it safe to take a heap snapshot in production?
With care. Writing a snapshot blocks the event loop and, according to the Node.js documentation, needs memory about twice the size of the heap. Take it from one instance removed from load balancing, write it to disk-backed storage, and treat the file as sensitive because it contains strings from memory, such as tokens and user data.
Should I increase --max-old-space-size to fix growing memory?
Not as a fix for a leak. A larger heap limit only delays the crash and makes each garbage collection pause longer. Raise it when measurements show the working set legitimately needs more heap, and keep it below the container memory limit so a real exhaustion fails with a JavaScript heap out of memory error instead of a silent kill.
What are the most common causes of Node.js memory leaks?
Unbounded caches in module-level Maps or objects, pending-request maps whose entries are never removed on timeout or error, event listeners added per request to long-lived emitters, intervals that are never cleared, metrics with unbounded label values, and Buffers retained by streams, slices, or ignored backpressure.
A Node.js memory leak is a structure that outlives the request that filled it. Tomosu maps where those structures are written and where they are cleaned up. Assess your repository →