Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Node.js memory

Node.js Memory Keeps Growing in Production: A Step-by-Step Investigation

Tomosu AI·14 min read·

The dashboard shows a Node.js service whose memory line only goes up. It restarts every few days, or the platform kills it, and the graph starts climbing again. Before anyone reaches for a bigger heap or a nightly restart, you need to answer two questions: is this actually a leak, and if so, which object in which code path is holding on to the memory?

Quick answer

When Node.js memory keeps growing in production, first check whether the heap floor after each full garbage collection is rising. A rising line on its own is not proof, because V8 lets garbage pile up between collections. If the floor rises under steady traffic, find the layer that grows, then compare heap snapshots to see what is retained.

This guide is a step-by-step investigation for a long-running Node.js service: how to read the growth curve, which process.memoryUsage() field matters, how to take and compare heap snapshots without causing a second incident, and the code patterns that show up at the end of most retainer chains. If your service runs in Kubernetes and the symptom is an OOMKilled pod, start with Kubernetes OOMKilled: Memory Leak, Memory Limit, or Something Else?, which covers heap caps, container limits, and what the cgroup counts. This post goes deeper into the Node.js side.

Is growing Node.js memory actually a leak?

Not necessarily. A memory leak in a garbage-collected runtime is memory that stays reachable after the program no longer needs it. The garbage collector cannot free it because something still points to it. That is different from memory that is simply not collected yet.

V8, the JavaScript engine inside Node.js, is lazy on purpose. Short-lived objects are collected cheaply and often in the young generation. Objects that survive move to the old generation, which is collected by a much more expensive full mark-compact pass. V8 runs that pass when it decides the heap has grown enough, not when your dashboard would like it to. Between full collections, heapUsed climbs with garbage that is already unreachable. heapTotal, the memory V8 has reserved for the heap, can also grow toward the heap limit and stay there.

So the raw memory line can go up for hours in a perfectly healthy process. The signal that matters is the floor: how much heap is still in use right after a full collection.

WATCH THE FLOOR AFTER EACH FULL GC, NOT THE PEAKS HEALTHY SAWTOOTH WARM-UP, THEN PLATEAU LEAK: FLOOR RATCHETS UP Floor after GC is flat. Not a leak. Floor rises, then levels off. Caches filling. Check bounds. Floor rises with traffic. Something retains objects. Solid line: heapUsed. Dashed line: heap still in use after a full garbage collection. Only the third shape is a leak, and it can take hours of traffic before it is visible.
All three services “keep growing” for the first hour. Only one of them has a leak.

The easiest way to see the floor is V8’s own GC trace. Start the process with --trace-gc (a V8 flag Node passes through) and read the full collections. It is verbose, so enable it on one instance for a limited period:

--trace-gc · full collections only (format varies by version)
# node --trace-gc server.js 2>&1 | grep -E 'Mark-(Compact|sweep)'
[1:0x5a1c]  3600512 ms: Mark-Compact 212.4 (231.0) -> 148.2 (229.5) MB, ...
[1:0x5a1c]  7201877 ms: Mark-Compact 239.8 (258.3) -> 171.9 (256.0) MB, ...
[1:0x5a1c] 10803004 ms: Mark-Compact 266.1 (284.7) -> 195.6 (282.1) MB, ...
# The number after "->" is heap in use after the collection. +24 MB per hour: a leak candidate.

Older Node.js versions label the full collection Mark-sweep instead of Mark-Compact. The number before the arrow is heap in use before the collection, the number after is what survived, and the values in parentheses are the reserved heap. If you already export metrics, the same trend is visible by plotting the minimum of heapUsed over a sliding window of several minutes.

Which memory number is growing: heap, external, or RSS?

A Node.js process has more memory than the V8 heap. process.memoryUsage() reports the layers separately, and knowing which one grows decides every next step. Heap snapshots, the main tool for JavaScript leaks, see only one of them.

FieldWhat it measuresIf it is the one growing
heapUsedLive JavaScript objects plus garbage not yet collectedWatch its post-GC floor. A rising floor means retained JavaScript objects. Heap snapshots will find them.
heapTotalHeap memory V8 has reservedGrowth toward the limit alone is normal. Read it together with heapUsed.
externalMemory of C++ objects bound to JavaScript objectsUsually Buffers (see next row). Otherwise native addons that allocate through V8.
arrayBuffersBacking stores of ArrayBuffer and Node.js Buffer (also counted in external)Retained Buffers: stream buffers, sliced Buffers, request bodies held in memory.
rssEverything resident for the process: heap, code, stacks, native allocationsRising while the others are flat means native memory or allocator fragmentation. Snapshots will not help.
WHICH LAYER IS GROWING? ASK IN ORDER Q1 · V8 HEAP Is the heapUsed floor after full GC rising? Q2 · OFF-HEAP BUFFERS Are external or arrayBuffers rising? Q3 · WHOLE PROCESS Is rss rising while heap and external are flat? JS objects retained Compare heap snapshots Buffers retained Streams, slices, backpressure Native or allocator Addons, fragmentation, stacks YESYESYES NONONO Probably not a leak The heap is growing toward its limit before a full GC, or traffic and payload sizes changed. Compare with load.
Heap snapshots answer only the first question. Decide which layer grows before you take one.

A short periodic log line with these fields, emitted every 15 to 60 seconds, is enough to answer all three questions. If you already use prom-client, its default metrics export heap and resident memory. Check that arrayBuffers or external is included, because that is the layer teams most often forget.

Step by step: how do you investigate a Node.js memory leak in production?

The procedure below works whether the service runs on a VM, in a container, or on a platform that restarts it for you. Steps 1 to 3 use data you probably already have. Steps 4 to 6 need a heap snapshot tool and one instance you can set aside.

  1. Confirm the shape of the growth. Plot the heap floor after full GCs over at least several hours, next to request rate. Flat floor: stop here. Rising then flat: check that caches have bounds. Rising steadily: continue.
  2. Find which layer is growing. Use the decision tree above. Only continue with heap snapshots if the heapUsed floor is the one rising.
  3. Correlate growth with a driver. Divide growth by request count to get bytes retained per request, and check whether growth follows requests, wall-clock time, unique keys, or a specific deploy (see the table below).
  4. Capture three heap snapshots. From one instance: after warm-up, after a period of normal load, and after a second period of load.
  5. Compare the snapshots. In the third snapshot, view the objects allocated between the first and second. They survived a whole load cycle and several full GCs.
  6. Follow the retainer chain to code. Pick the largest surviving group and read its retainers until you reach a variable that your code owns.
  7. Fix and verify under the same load. The fix is proven only when the post-GC floor stays flat under the traffic pattern that produced the growth.

Step 3 in practice: what does the growth follow?

This step is cheap and often skipped. It narrows the search before you open a single snapshot. Suppose the floor grows 24 MB per hour while the instance serves 40 requests per second, which is 144,000 requests per hour. That is roughly 170 bytes per request. That is too small for a leaked request body, but it fits a map entry keyed by request ID, or one listener closure per request.

Growth followsSuspect first
Request count, including at nightSomething stored per request: a pending-request map, a per-request listener, a per-request closure in a long-lived array.
Wall-clock time, even with no trafficA timer: setInterval that appends to a structure, or intervals created and never cleared.
Number of distinct users, tenants, or URLsA keyed cache or metric without a bound: a memoization map, a per-tenant client, a metric label with user IDs.
Specific endpoint or payload sizeBuffers or parsed payloads held by one handler; stream backpressure ignored on large downloads.
A particular deployDiff the release. The commit investigation workflow applies to memory regressions too.

Steps 4 and 5: the three-snapshot comparison

A single heap snapshot tells you what is big. It does not tell you what is growing, and big caches that are working as designed dominate it. Comparing snapshots removes that noise. The three-snapshot technique goes one step further: it isolates objects that were created during one load period and were still alive a full load period later.

THE THREE-SNAPSHOT TECHNIQUE Load period A objects created here Load period B several full GCs run S1 S2 S3 after warm-up after load A after load B DROPS OUT OF THE VIEW Request objects created in A and freed. Caches that were already full at S1. IN S3: ALLOCATED BETWEEN S1 AND S2 Created in A, still alive after all of B. These survivors are the leak candidates. Each snapshot forces a full GC first, so garbage never appears in the comparison.
Objects that were created in load period A and are still alive after load period B are what you are looking for.

To take the snapshots from a running process, start it with --heapsnapshot-signal=SIGUSR2 and send that signal, or call v8.writeHeapSnapshot() from a protected admin route. Load the three .heapsnapshot files in the Chrome DevTools Memory panel. Then:

The DevTools heap snapshot documentation describes each view, and the Node.js guide Using Heap Snapshot covers the capture options.

How do you read a retainer chain?

A retainer chain is the path of references from a GC root to the object you selected. The Retainers pane at the bottom of the Memory panel shows it from the object upward. Everything in that chain is a reason the object cannot be collected. Your job is to find the first link that belongs to your code and should not be there.

READ THE CHAIN UNTIL YOU REACH YOUR OWN CODE (ILLUSTRATIVE) (GC roots) module scope of payment-client.js pending: Map (48,210 entries) entry { resolve, reject, timer } closure context: req, body Buffer Always reachable. Not interesting. Module-level variables live forever. Fix point: the first link you own that grows without a bound. One entry per outstanding request. What the entry drags along with it. The large retained size shows on the Map. The fix is to delete entries on timeout and error too.
The biggest object is rarely the bug. The bug is the long-lived container that keeps adding references to it.

Three reading rules save most of the time:

Which code patterns cause most Node.js memory leaks?

Almost every retainer chain ends in one of a handful of patterns. All of them share one property: a structure that lives for the life of the process gains an entry per request, per key, or per tick, and the code that removes the entry runs on only some paths.

PatternWhat the snapshot showsFix
Pending-request mapA Map of resolvers or callbacks whose count keeps risingDelete the entry in a finally, on timeout, and on error, not only on success
Unbounded cache or memoizationA module-level Map or object keyed by user, URL, or queryA bounded cache (lru-cache with max and ttl), or no cache
Listener added per requestGrowing listener arrays on a long-lived emitter; MaxListenersExceededWarningRemove the listener when the request ends, or use { once: true }
Timers never clearedGrowing Timeout objects and the closures they captureKeep the handle; clearInterval/clearTimeout on every exit path
Metric label cardinalityMetrics registry objects growing with user IDs or raw pathsLabel by route template and status, never by ID
Retained BuffersGrowing arrayBuffers; small slices keeping large parent Buffers aliveBuffer.from(slice) to copy what you keep; respect backpressure with pipeline()
In-memory session storeSession objects accumulating per visitorAn external session store; the default express-session MemoryStore warns it is not for production

The pending-request map

This is the pattern in the retainer chain above, and it is common in hand-written RPC clients, WebSocket request/response layers, and queue consumers that wait for replies. The success path cleans up. The timeout and error paths do not.

payment-client.jsleaks on timeout
const pending = new Map();            // lives as long as the process

function call(msg) {
  return new Promise((resolve, reject) => {
    pending.set(msg.id, { resolve, reject });
    setTimeout(() => reject(new Error('timeout')), 5000);  // entry is never deleted
    socket.send(JSON.stringify(msg));
  });
}

socket.on('message', (raw) => {
  const reply = JSON.parse(raw);
  const entry = pending.get(reply.id);
  if (entry) { pending.delete(reply.id); entry.resolve(reply); }  // only the success path cleans up
});
payment-client.jscleaned up on every path
function call(msg, timeoutMs = 5000) {
  return new Promise((resolve, reject) => {
    const timer = setTimeout(() => {
      pending.delete(msg.id);                          // timeout path
      reject(new Error('timeout'));
    }, timeoutMs);
    pending.set(msg.id, {
      resolve: (v) => { clearTimeout(timer); pending.delete(msg.id); resolve(v); },
      reject:  (e) => { clearTimeout(timer); pending.delete(msg.id); reject(e); },
    });
    try { socket.send(JSON.stringify(msg)); }
    catch (err) { pending.get(msg.id)?.reject(err); }    // send failure path
  });
}

Note what is not the leak here. A promise that never settles is not retained by itself; if nothing references it, it is collected with its callbacks. The leak is the long-lived Map that holds the resolver, and through its closure the caller’s context.

The listener added per request

Node warns when an emitter gets more than 10 listeners for one event, which is the default maxListeners. The warning names the event and the emitter type:

stderrsymptom
(node:1) MaxListenersExceededWarning: Possible EventEmitter memory leak detected.
  11 change listeners added to [EventEmitter]. Use emitter.setMaxListeners() to increase limit
(node:1) MaxListenersExceededWarning: Possible EventTarget memory leak detected.
  11 abort listeners added to [AbortSignal]. Use events.setMaxListeners() to increase limit

The fix is almost never setMaxListeners(), which only silences the warning. Run with --trace-warnings to get the stack of the line that added the eleventh listener, then remove the listener when the request is done:

handler.jslistener removed when the request ends
// Broken: one listener per request on an emitter that lives forever
configBus.on('change', () => res.locals.stale = true);

// Fixed: register once at startup, or remove it when the request finishes
const onChange = () => { res.locals.stale = true; };
configBus.on('change', onChange);
res.on('close', () => configBus.off('change', onChange));   // released per request

The unbounded cache

An in-process cache is a leak with good intentions. It looks like the warm-up curve in the first diagram until the key space turns out to be much larger than expected: every user, every URL with a query string, every tenant. Give every cache a size bound and an expiry, and decide what happens when it is full.

pricing.jsbounded by count and age
// Broken: grows with every distinct (sku, region, currency) ever requested
const cache = {};

// Fixed: lru-cache v10+
import { LRUCache } from 'lru-cache';
const cache = new LRUCache({ max: 10_000, ttl: 5 * 60 * 1000 });
Metrics can leak too

A Prometheus client keeps one time series in memory for every distinct combination of label values. A histogram labelled with userId or the raw request path (/orders/8812) grows with every new value. The retainer chain ends in the metrics registry, not in your business code. Label by route template (/orders/:id) instead.

When is growing memory not a leak?

Some of the most expensive memory investigations end with “nothing is leaking.” Rule these out before you blame the code:

Restarts hide the evidence

A scheduled restart or an aggressive memory-based restart policy keeps the service up, but it also resets the curve before the leak becomes obvious. If you use one as a stopgap, keep one instance running long enough to capture the three snapshots.

How do you capture evidence safely in production?

Heap snapshots are the best evidence and the most disruptive to collect. According to the Node.js v8 documentation, writing a snapshot is synchronous, so it blocks the event loop. It also needs memory about twice the size of the heap at that moment. On a process that is already close to its limit, taking the snapshot can cause the crash you are investigating.

MethodCostUse it when
process.memoryUsage() metricsNegligibleAlways. It answers the “which layer” question.
--trace-gcLog volumeFor a limited time on one instance to read the post-GC floor.
--heap-prof (sampling heap profiler)Low; written when the process exitsTo see which call sites allocate the most over a run, in a test or canary.
Heap snapshot (--heapsnapshot-signal, v8.writeHeapSnapshot())Blocks the event loop; ~2× heap memoryOn one instance removed from the load balancer, early in the growth curve.
--heapsnapshot-near-heap-limit=NSame, at the worst momentAs a last resort to catch the state just before a heap out-of-memory crash.
Inspector attached via --inspectInteractive, pauses on snapshotStaging, or production via an SSH tunnel. Never bind the inspector to a public interface.

Most Node.js memory leaks are cleanup code that runs on the success path and nowhere else.

Fixes, matched to the cause

CauseFixWhat not to do
Per-request entries in a long-lived structureRemove the entry on every exit path: success, error, timeout, client disconnect. Use finally.Add a periodic sweep that hides a missing cleanup path.
Unbounded cacheBound by count and age; measure hit rate; consider an external cache for large key spaces.Raise --max-old-space-size so the cache has more room.
Listener or timer leakRegister once, or pair every on/setInterval with an off/clearInterval tied to the owner’s lifetime.setMaxListeners(0) to silence the warning.
Retained BuffersStream instead of buffering; use stream.pipeline(); copy small slices you keep.Read whole files or bodies into memory “because they are small.”
Native memory or fragmentationIsolate the addon or allocation pattern; test allocator settings with measurements.Take more heap snapshots. They cannot see this memory.
Heap legitimately too smallRaise the heap limit with headroom, below any container limit.Treat this as the default fix before the floor is measured.

Verify with the same measurement that raised the alarm: run the traffic pattern that caused growth and watch the post-GC floor. A fix that makes the raw line look better for an hour proves nothing. If the leak arrived with a recent release, the same investigation pairs well with correlating logs, traces, and a code change. The same “acquire without release” shape appears for other resources in Resource Leaks in Java Services.

How Tomosu helps

A heap snapshot finds the structure that leaked today, in one process. It does not list the other structures written the same way that have not leaked yet. Tomosu analyzes the repository and each pull request for that list:

These findings roll up into the Production Reliability Index, so a change that adds a per-request entry to a process-wide map can be discussed in review rather than discovered from a memory graph.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

Is it normal for Node.js memory to keep growing?

Some growth is normal. V8 collects garbage lazily, so heapUsed rises between collections and heapTotal can grow toward the heap limit before a full collection runs. What is not normal is a heap floor, measured right after full garbage collections, that keeps rising for hours under steady traffic. That pattern points to a leak.

How do I find a memory leak in a Node.js application in production?

Confirm the post-GC heap floor is rising, identify whether heapUsed, arrayBuffers, or only rss is growing, then take three heap snapshots from one instance at intervals under load. In the third snapshot, look at objects allocated between the first and second, and follow their retainers back to the Map, listener, timer, or closure in your code that holds them.

Why does RSS keep growing when the Node.js heap is flat?

RSS counts all memory the process has resident, not just the V8 heap. If heapUsed and external are flat while RSS rises, suspect native memory from addons, memory held by the allocator after fragmentation, or thread stacks. Heap snapshots will not show this memory, so compare process.memoryUsage() fields before spending time on snapshots.

Is MaxListenersExceededWarning a memory leak?

It is a strong hint, not proof. Node emits it when more than 10 listeners for one event are added to one emitter by default. When the count keeps climbing, code is usually adding a listener per request to a long-lived emitter and never removing it. Run with --trace-warnings to get the stack trace of the code that added the listener.

Is it safe to take a heap snapshot in production?

With care. Writing a snapshot blocks the event loop and, according to the Node.js documentation, needs memory about twice the size of the heap. Take it from one instance removed from load balancing, write it to disk-backed storage, and treat the file as sensitive because it contains strings from memory, such as tokens and user data.

Should I increase --max-old-space-size to fix growing memory?

Not as a fix for a leak. A larger heap limit only delays the crash and makes each garbage collection pause longer. Raise it when measurements show the working set legitimately needs more heap, and keep it below the container memory limit so a real exhaustion fails with a JavaScript heap out of memory error instead of a silent kill.

What are the most common causes of Node.js memory leaks?

Unbounded caches in module-level Maps or objects, pending-request maps whose entries are never removed on timeout or error, event listeners added per request to long-lived emitters, intervals that are never cleared, metrics with unbounded label values, and Buffers retained by streams, slices, or ignored backpressure.


A Node.js memory leak is a structure that outlives the request that filled it. Tomosu maps where those structures are written and where they are cleaned up. Assess your repository →