Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Kubernetes memory

Kubernetes OOMKilled: Memory Leak, Memory Limit, or Something Else?

Tomosu AI·16 min read·

A Kubernetes OOMKilled status tells you the kernel killed a container because it reached its memory limit. It does not tell you why the container needed that memory. Before you raise the limit, you need to know whether the process is leaking, spiking on specific work, using memory outside the heap your dashboard shows, or simply running with a limit that was never big enough.

Quick answer

OOMKilled (exit code 137) means the Linux kernel sent SIGKILL to a container whose cgroup memory usage reached its resources.limits.memory and could not be reclaimed. The limit is where the failure happened, not necessarily the cause. Common causes:

The rest of this guide shows how to tell those apart with evidence: the container’s termination state, cgroup counters, the right memory metric, heap versus RSS, and a map of the code paths that accumulate memory. Examples use Node.js, with JVM equivalents where they differ.

What does OOMKilled (exit code 137) mean in Kubernetes?

Every container with a memory limit runs inside a Linux cgroup whose ceiling is set from resources.limits.memory (memory.max on cgroup v2, memory.limit_in_bytes on v1). When the processes in that cgroup try to allocate past the ceiling, the kernel first tries to reclaim memory, mostly clean file cache. If it cannot free enough, the cgroup OOM killer picks a process in that cgroup and sends it SIGKILL. The kubelet notices the container died from an OOM kill and records it:

kubectl describe pod api-7c9f8d6b5-x2k4qsymptom
    State:          Running
      Started:      Tue, 29 Sep 2026 10:42:11 +0000
    Last State:     Terminated
      Reason:       OOMKilled
      Exit Code:    137
      Started:      Tue, 29 Sep 2026 06:15:02 +0000
      Finished:     Tue, 29 Sep 2026 10:42:09 +0000
    Restart Count:  14
    Limits:
      memory:  512Mi
    Requests:
      memory:  256Mi

Exit code 137 is 128 + 9: the process was terminated by signal 9, SIGKILL. Because SIGKILL cannot be caught, the application gets no chance to log anything. That is why an OOM kill usually leaves the application logs looking normal right up to the last line.

137 is not always an OOM kill

Any SIGKILL produces exit code 137. A container that ignores SIGTERM during a failed liveness probe or a rolling update is killed after terminationGracePeriodSeconds and also exits 137, but its reason is Error, not OOMKilled. Trust the reason field, not the exit code alone.

On the node, the kernel logs every OOM kill. The constraint field tells you whether a container limit (CONSTRAINT_MEMCG) or the whole node ran out (CONSTRAINT_NONE):

Node · journalctl -k or dmesg
oom-kill:constraint=CONSTRAINT_MEMCG,nodemask=(null),cpuset=...,oom_memcg=/kubepods/burstable/pod...,task=node,pid=48213,uid=1000
Memory cgroup out of memory: Killed process 48213 (node) total-vm:1893420kB,
  anon-rss:515204kB, file-rss:31120kB, shmem-rss:0kB, UID:1000 pgtables:1680kB oom_score_adj:984

The anon-rss and shmem-rss values are the first clue to what filled the limit. Here almost all of the 512 Mi was anonymous memory: heap, Buffers, or native allocations, not files.

OOMKilled, Evicted, or heap out of memory: which one happened?

Three different mechanisms end a container because of memory, and teams regularly confuse them. They are enforced by different components, they leave different evidence, and they need different fixes.

THREE WAYS MEMORY ENDS A CONTAINER WHO ACTSWHO ACTSWHO ACTS Linux kernel (cgroup)kubeletLanguage runtime TRIGGERTRIGGERTRIGGER Container usage reachesits own limit (memory.max) Node memory.availablefalls below the threshold Heap reaches its own cap(--max-old-space-size) WHAT YOU SEEWHAT YOU SEEWHAT YOU SEE reason: OOMKilledexit code 137 (SIGKILL) status: Failedreason: Evicted heap out of memoryexit 134 (Node.js) No error in app logs “node was low on memory” reason: Error, stack in logs Only the first is Kubernetes OOMKilled. Each needs a different fix, so confirm which one fired.
The same incident title, “pod died of memory”, covers three mechanisms with three different owners.

Node-pressure eviction

When the node as a whole runs low, the kubelet evicts pods before the kernel has to step in. It watches memory.available (node capacity minus the node’s working set) against eviction thresholds; the default hard threshold on Linux is memory.available<100Mi. Evicted pods end in phase Failed with reason Evicted, and the kubelet ranks candidates by whether their usage exceeds their requests, then by priority, then by how far above requests they are. A pod can be evicted without ever reaching its own limit. The fix is usually honest memory requests, not a higher limit. See the Kubernetes docs on node-pressure eviction.

If memory runs out faster than the kubelet can evict, the kernel’s system-wide OOM killer acts. It picks victims using oom_score_adj, which the kubelet sets by QoS class: -997 for Guaranteed, 1000 for BestEffort, and a value in between for Burstable that depends on the memory request. That kill is also reported as OOMKilled, but the kernel log shows CONSTRAINT_NONE and the node usually shows memory pressure at the same time.

Runtime heap limit

Node.js and the JVM enforce their own heap caps inside the process. When V8 cannot grow the heap past --max-old-space-size, Node prints FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory and aborts. The container exits with 134 and reason Error. This is the better failure: you get a message, a stack, and optionally a heap snapshot. It also means the heap cap, not the container limit, was the binding constraint.

Shell · confirm which mechanism fired
# Termination reason and exit code for every container, including sidecars
kubectl get pod api-7c9f8d6b5-x2k4q -o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.lastState.terminated.reason}{"\t"}{.lastState.terminated.exitCode}{"\n"}{end}'

# Evictions and node memory pressure
kubectl get events -A --field-selector reason=Evicted
kubectl describe node <node> | grep -A2 MemoryPressure

# Inside a running container on cgroup v2: limit, current usage, peak, and OOM counters
cat /sys/fs/cgroup/memory.max /sys/fs/cgroup/memory.current /sys/fs/cgroup/memory.peak
cat /sys/fs/cgroup/memory.events      # high, max, oom, oom_kill counts
grep -E '^(anon|file|shmem|kernel|sock|active_file|inactive_file) ' /sys/fs/cgroup/memory.stat

On cgroup v1 the equivalents are memory.limit_in_bytes, memory.usage_in_bytes, memory.max_usage_in_bytes, and memory.failcnt. memory.peak needs a recent kernel (5.19 or later).

What counts against a Kubernetes container memory limit?

Everything the container’s processes cause the kernel to charge to its cgroup counts. That is more than the heap your application dashboard shows, and more than most people expect:

WHAT THE CGROUP CHARGES, AND WHAT EACH METRIC SEES limits.memory = memory.max V8 / JVM heap Buffers,off-heap native,stacks tmpfs,emptyDir activefile cache inactivefile cache heapUsed (app dashboards) container_memory_rss (anonymous memory) container_memory_working_set_bytes (kubectl top, eviction) container_memory_usage_bytes (all charged memory) The kernel reclaims clean file cache before it kills anything. Heap, Buffers, native memory, and tmpfs cannot be reclaimed without swap, so they drive OOM kills.
A heap chart can sit at 40% of the limit while anonymous memory and tmpfs fill the rest.

Two metric details matter during an investigation. First, container_memory_working_set_bytes is usage minus inactive_file. It is what kubectl top and the eviction manager use, and it is the best single proxy for “how close is this container to being killed”. It still includes active file cache, which the kernel can reclaim, so a working set near the limit does not guarantee a kill is coming. Second, container_memory_rss counts anonymous memory only. It excludes tmpfs, which is accounted as shared memory (shmem in memory.stat). If working set is high and RSS is not, look at tmpfs and file cache.

Why is my pod OOMKilled but memory looks low?

This is the most common question in Stack Overflow threads about OOMKilled, and there are five usual explanations. Check them in this order.

  1. You are looking at the wrong metric. A heap graph, process.memoryUsage().heapUsed, or JVM “heap used” is a fraction of what the cgroup charges. Compare working set to the limit, per container.
  2. The spike fell between scrapes. Prometheus typically scrapes every 15 to 60 seconds. A request that allocates 400 MB and gets killed within two seconds never appears in a stored sample.
  3. You are looking at the wrong container. Pod-level graphs sum or average containers. The killed container may be a service mesh sidecar, a log shipper, or an init container with its own small limit.
  4. Aggregation hid it. avg across replicas, or avg_over_time over five minutes, flattens the one pod that hit the limit. Use max.
  5. The memory is not in the process. tmpfs files in an emptyDir with medium: Memory are charged to the container even though no process shows them in its RSS.
WHAT A 30 SECOND SCRAPE INTERVAL STORES limit 0 OOM kill about 2 s after the request started restart Stored samples: peak about 25% Actual peak: 100% of the limit 0 s30 s60 s90 s120 s Red: actual cgroup usage. Purple: the samples Prometheus stored at a 30 s scrape interval. The spike and the kill both fall between two scrapes, so no stored sample is near the limit.
Sampled metrics are good at slow growth and bad at fast spikes. The kill counter and the kernel log do not miss either.

For spikes, rely on counters rather than gauges. The restart count and last termination reason from kube-state-metrics are exact, and memory.peak or memory.events inside the container record the high-water mark and every OOM event regardless of scrape timing. Useful queries:

PromQL
# Working set as a fraction of the limit, per container (max, not avg)
max by (namespace, pod, container) (container_memory_working_set_bytes{container!="", container!="POD"})
  / on (namespace, pod, container)
max by (namespace, pod, container) (kube_pod_container_resource_limits{resource="memory"})

# Containers whose last termination was an OOM kill and that restarted recently
kube_pod_container_status_last_terminated_reason{reason="OOMKilled"} == 1
  and on (namespace, pod, container)
increase(kube_pod_container_status_restarts_total[1h]) > 0

# Heap vs working set gap (needs an app-level heap metric, e.g. prom-client defaults)
max by (pod) (container_memory_working_set_bytes{container="api"})
  - on (pod) max by (pod) (nodejs_heap_size_used_bytes)

Kubernetes OOMKilled causes: memory leak, load spike, or limit too low?

Once you have confirmed a genuine OOM kill in your own container, the question is what shape the memory had before it died. The four shapes below account for most incidents, and each points at a different fix.

FOUR SHAPES BEFORE AN OOM KILL LEAKLOAD SPIKELIMIT TOO LOWOUTSIDE THE HEAP heap rss Climbs with uptime;a restart resets it Flat, then a jumpon one request or job Near limit at start;any bump kills it Heap flat, RSS climbsBuffers/native/tmpfs Plot working set against the limit over days, and overlay restarts and deploys, before choosing a fix.
The shape over time separates a leak from a spike from an undersized limit. A single snapshot cannot.

1. Leak or unbounded growth

Memory rises with uptime, independent of traffic, and never returns to its starting level. Restart intervals are roughly regular, and a deploy “fixes” it for a while. Strictly, garbage-collected languages rarely leak in the C sense. What they have is unbounded retention: a structure that is still reachable and keeps growing, such as a module-level cache, a listener list, or a map of sessions nobody removes.

2. Load-driven spike

Baseline is healthy, and kills line up with a particular endpoint, batch job, tenant, or input size. The code is usually buffering a whole payload: reading a full query result, building a whole CSV or JSON document in memory, decompressing an upload, or running unbounded Promise.all over a large list. Memory per request × concurrent requests exceeds the limit.

3. Limit too low for the steady state

Working set climbs during warm-up to a plateau close to the limit and stays there. Nothing grows over time; there is just no headroom. This is the only cause where raising the limit is the fix, and it is also common when the heap cap was set larger than the container, so the runtime happily fills memory the cgroup does not have.

4. Memory outside the heap

The heap graph is flat or small, while RSS or working set climbs. Suspects: Buffers retained by a stream that is never consumed, a native module, allocator fragmentation, unclosed compression streams, JVM direct buffers or thread growth, and files written to tmpfs.

The decision tree below puts the checks in the order that eliminates the most confusion first.

WHICH CAUSE IS IT? SIX QUESTIONS, IN ORDER Q1 · POD STATUSQ2 · EXIT CODE AND LOGSQ3 · WHICH CONTAINER Q4 · WORKING SET OVER DAYSQ5 · CORRELATIONQ6 · HEAP VS RSS Is the pod status reason Evicted, not OOMKilled? Exit 134 and “heap out of memory” in the logs? Was the killed container a sidecar, not your app? Does working set climb with uptime, even at low load? Do kills line up with one endpoint, job, or payload? Is the heap small while RSS or working set is large? Node-pressure eviction Runtime heap limit Another container Leak or unbounded growth Load-driven spike Memory outside the heap Node low; pods above requests Heap cap hit before cgroup Sidecar or init container Find the accumulation path Bound per-request memory Buffers, native, tmpfs YESYESYESYESYESYES NONONONONONO Limit too low for the steady state Usage is stable and legitimate. Raise the limit from the measured peak plus headroom.
Raising the memory limit is the right answer only when all six questions come back “no”.
SignalLeakLoad spikeOutside the heapLimit too low
Working set over daysRamps with uptime, sawtooth at each restartFlat, with sharp jumpsRamps or stepsPlateau near the limit
After traffic dropsStays highReturns to baselineOften stays highStays at plateau
Heap vs RSSHeap grows with RSSHeap spikes if JS objects; RSS if BuffersHeap flat, RSS or shmem growsBoth stable
Correlates withUptime, distinct keys or users seenAn endpoint, job, tenant, or input sizeStreams, uploads, native modules, tmpfs writesEvery pod, from startup
Heap snapshotsRetained count of one type grows between snapshotsLarge transient arrays or stringsLittle changeLittle change
Kernel loganon-rss near limitanon-rss near limitshmem-rss or anon outside heapanon-rss near limit

How do you find a Node.js memory leak in production?

A Node.js memory leak in production is hard mostly because the evidence lives in three layers: V8’s heap, the process’s memory outside the heap, and the cgroup. Start by making the layers visible, then take heap snapshots only when you know the heap is the layer growing.

Step 1: Set the heap cap below the container limit

--max-old-space-size sets the V8 old-generation heap limit in megabytes. The default is derived from the memory Node detects, and depending on version it may not match your container limit. Set it explicitly and leave room for everything that is not heap: Buffers, native memory, stacks, and code.

deployment.yaml · heap cap below the cgroup limitfails with a stack, not SIGKILL
containers:
  - name: api
    image: registry.example.com/api:1.42.0
    env:
      - name: NODE_OPTIONS
        value: "--max-old-space-size=640 --heapsnapshot-signal=SIGUSR2"
    resources:
      requests:
        memory: "768Mi"
      limits:
        memory: "1Gi"   # ~384 MiB left for Buffers, native, stacks, code

With the cap below the limit, a heap leak ends as a JavaScript heap out-of-memory crash with a message and a stack. With the cap above the limit, the same leak ends as a silent OOM kill. How much headroom to leave depends on how much the service holds outside the heap, which is the next thing to measure.

Step 2: Export heap and non-heap memory separately

memory-report.js
const MB = 1024 * 1024;
setInterval(() => {
  const m = process.memoryUsage();
  logger.info({
    rss_mb:           Math.round(m.rss / MB),           // resident set: what the kernel sees
    heap_total_mb:    Math.round(m.heapTotal / MB),     // V8 heap reserved
    heap_used_mb:     Math.round(m.heapUsed / MB),      // live JS objects + garbage not yet collected
    external_mb:      Math.round(m.external / MB),      // C++ objects bound to JS, incl. Buffers
    array_buffers_mb: Math.round(m.arrayBuffers / MB),  // Buffer / ArrayBuffer backing stores
  }, 'memory');
}, 15_000).unref();

If you use prom-client, its default metrics already export heap and RSS figures. Read them together: heapUsed growing points at retained JavaScript objects. external or arrayBuffers growing points at Buffers. RSS growing while all of those are flat points at native memory or allocator fragmentation.

Step 3: Compare heap snapshots

A single heap snapshot shows what is big. Two or three snapshots taken minutes or hours apart show what is growing, which is what you need for a leak. With --heapsnapshot-signal=SIGUSR2 set, send the signal to the Node process to write a .heapsnapshot file to the working directory. Alternatively call v8.writeHeapSnapshot() from an admin endpoint, or use --heapsnapshot-near-heap-limit=N to capture automatically as the heap approaches its cap.

Shell · snapshot from one pod
# Take the pod out of rotation first if you can: the process pauses while writing
kubectl exec api-7c9f8d6b5-x2k4q -- sh -c 'kill -USR2 $(pgrep -o node)'
kubectl exec api-7c9f8d6b5-x2k4q -- ls -la /app/*.heapsnapshot
kubectl cp api-7c9f8d6b5-x2k4q:/app/Heap.20260929.104512.1.0.001.heapsnapshot ./snap-1.heapsnapshot

# Repeat later, then load both in Chrome DevTools > Memory and use the Comparison view
A heap snapshot can cause the OOM kill you are investigating

The Node.js v8 documentation notes that creating a heap snapshot requires memory about twice the size of the heap at the time. A process using 600 MB of heap in a 1 GiB container may be killed mid-snapshot. Capture on a pod with a temporarily raised limit, or capture early in the growth curve and compare, and make sure the snapshot is written to a disk-backed volume, not tmpfs.

In the Comparison view, sort by Size Delta or # Delta and open the retainers of the growing type. The retainer chain usually ends at a module-level variable: a Map, an array, an emitter’s listener list, or a closure captured by a timer. That variable is the accumulation path.

If snapshots show no heap growth while RSS climbs, stop looking at JavaScript objects. Check arrayBuffers for Buffers retained by streams, native addons, and whether the service writes to a memory-backed volume.

The same questions for the JVM

The JVM has the same layers with different names, and the same trap: the heap is only part of what the cgroup charges.

QuestionNode.jsJVM
Heap cap--max-old-space-size-Xmx, or -XX:MaxRAMPercentage (default 25% of the container limit on container-aware JDKs)
Heap exhausted“JavaScript heap out of memory”, process abortsOutOfMemoryError: Java heap space; add -XX:+ExitOnOutOfMemoryError so the pod restarts cleanly
Outside the heapBuffers, native addons, stacksMetaspace, thread stacks, code cache, direct buffers (-XX:MaxDirectMemorySize), GC structures
Non-heap breakdownprocess.memoryUsage()-XX:NativeMemoryTracking=summary, then jcmd <pid> VM.native_memory summary
Heap dumpHeap snapshot, DevTools Comparisonjcmd <pid> GC.heap_dump, -XX:+HeapDumpOnOutOfMemoryError

HeapDumpOnOutOfMemoryError fires only when the JVM itself throws OutOfMemoryError. A cgroup OOM kill is SIGKILL, so the JVM never gets to write anything. See the Oracle Native Memory Tracking guide.

For Go services the equivalent lever is GOMEMLIMIT, a soft limit that makes the garbage collector work harder as the process approaches it. Set it below the container limit for the same reason.

Find the code paths that accumulate memory

Runtime evidence tells you that memory grows and, with snapshots, which type grows in one process. It does not give you the list of code paths that can grow without a bound, including the ones that have not caused an incident yet. That list comes from reading the code for a specific question: for every structure that lives longer than a request, what limits its size?

Accumulation pathWhat growsBounded byFix
Module-level Map or object used as a cacheOne entry per distinct keyNothing: distinct users, IDs, or URLsLRU with max and TTL, or an external cache
Listener added per request or connectionListener closures and everything they captureNothing until removedRemove on close, use once, or an AbortSignal
Whole-payload bufferingFull result sets, documents, uploadsInput size × concurrencyStream with backpressure; cap body sizes
Unbounded concurrencyIn-flight work for every itemLength of the input listConcurrency limit or batching
In-memory queue or retry bufferBacklog while a dependency is slowDownstream latencyBounded queue, shed or reject when full
Metrics labels with user or request IDsTime series in the client registryDistinct label valuesBounded label sets only
Files on emptyDir medium: Memorytmpfs pages charged to the containersizeLimit, if setDisk-backed volume, cleanup, sizeLimit

Unbounded cache keyed by request data

profile-cache.jsgrows with every distinct id
const cache = new Map();

export async function getProfile(id) {
  if (cache.has(id)) return cache.get(id);
  const profile = await db.profiles.findById(id);
  cache.set(id, profile);   // never evicted: size = number of distinct users since start
  return profile;
}
profile-cache.jsbounded by count and age
import { LRUCache } from 'lru-cache';

const cache = new LRUCache({ max: 10_000, ttl: 5 * 60_000 });  // at most 10k entries, 5 min each

export async function getProfile(id) {
  const hit = cache.get(id);
  if (hit) return hit;
  const profile = await db.profiles.findById(id);
  cache.set(id, profile);
  return profile;
}

This is the classic leak shape: in a load test with 50 test users, the map stops at 50 entries and looks fine. In production it grows with the number of distinct users since the last restart.

Listener registered per request

events-stream.jsbefore and after
// Before: every SSE client adds a listener that is never removed.
app.get('/events', (req, res) => {
  res.setHeader('Content-Type', 'text/event-stream');
  const onUpdate = (msg) => res.write(`data: ${JSON.stringify(msg)}\n\n`);
  bus.on('update', onUpdate);   // retains res and its buffers after the client leaves
});

// After: remove the listener when the connection closes.
app.get('/events', (req, res) => {
  res.setHeader('Content-Type', 'text/event-stream');
  const onUpdate = (msg) => res.write(`data: ${JSON.stringify(msg)}\n\n`);
  bus.on('update', onUpdate);
  req.on('close', () => bus.off('update', onUpdate));   // bounded by open connections
});

Node warns about this pattern with MaxListenersExceededWarning: Possible EventEmitter memory leak detected once an emitter has more than 10 listeners for one event. Teams often silence it with setMaxListeners(0). Treat that call in a diff as a question, not a fix.

Buffering a whole payload

export.jsmemory = rows × concurrency
app.get('/export', async (req, res) => {
  const { rows } = await pool.query(EXPORT_SQL, [req.query.since]);  // entire result set
  const csv = rows.map(toCsvLine).join('\n');                        // a second full copy
  res.type('text/csv').send(csv);
});
export.jsmemory bounded by batch size
import QueryStream from 'pg-query-stream';
import { pipeline } from 'node:stream/promises';

app.get('/export', async (req, res) => {
  const client = await pool.connect();
  try {
    const rows = client.query(new QueryStream(EXPORT_SQL, [req.query.since], { batchSize: 500 }));
    res.type('text/csv');
    await pipeline(rows, toCsv(), res);   // toCsv(): object-mode Transform; backpressure end to end
  } finally {
    client.release();
  }
});

The first version works in staging, where the export returns 2,000 rows. It becomes a load-driven spike the first time a large tenant exports two years of data, and it multiplies when two such exports run at once. Upload handlers have the same shape: a body parser configured with a very large limit, or await file.arrayBuffer() on an unbounded upload.

Fixes, matched to the cause

CauseFixWhat not to do
Leak or unbounded growthFind the retaining structure from snapshot comparison, then bound it: size limits, TTLs, listener removal. Find the sibling paths written the same way.Raise the limit. A leak fills any limit, only later.
Load-driven spikeStream instead of buffer, cap request and upload sizes, limit concurrency per pod.Size every pod for the largest tenant’s worst request.
Memory outside the heapMeasure the layer (arrayBuffers, NMT, shmem), then fix streams, native usage, or tmpfs writes.Tune the heap flag. It does not govern this memory.
Runtime heap limitTreat as a leak or spike investigation with a free stack trace. Raise the cap only if the heap plateau is legitimate and the container has room.Set the heap cap above the container limit.
Node-pressure evictionSet memory requests from measured usage; use Guaranteed QoS for critical pods; check node allocatable.Raise limits without raising requests.
Limit too lowRaise the limit from the measured peak (memory.peak, max working set) plus headroom, and keep the heap cap consistent with it.Remove the limit. The node then absorbs the problem.

A bigger memory limit turns a leak that restarts every four hours into one that restarts every eight.

Evidence to require before any fix

Memory limit bumps are the easiest change to approve in a review: one line of YAML, obviously safe, and the alerts stop for a while. Whether the proposal comes from a teammate, an answer online, or an AI coding assistant, ask for the evidence that matches it. This is the same discipline as the connection pool timeouts in part one of this series: the error names the resource that ran out, not the code that used it up.

Before raising the limit
  • Termination reason is OOMKilled for the app container
  • Working set plateaus; it does not ramp with uptime
  • No correlation with one endpoint or input size
  • Heap cap is consistent with the new limit
Before calling it a leak
  • Growth continues at low traffic
  • Snapshot comparison shows a growing retained type
  • The retaining structure is named in code
  • What bounds it, or why nothing does, is shown
Before blaming the platform
  • Kernel log shows CONSTRAINT_NONE or eviction events
  • Node MemoryPressure at the time
  • Requests compared with real usage on that node

For how this fits into release decisions more broadly, see Pre Merge Reliability Analysis and Production Reliability vs Observability. Observability tells you a pod died; reliability analysis asks which code made it likely.

How Tomosu helps

Runtime data and repository review answer different halves of the OOMKilled question. Metrics, the kernel log, and heap snapshots confirm that memory grows and which layer grows. A review of the repository finds which code paths can grow without a bound, including those that have not caused an incident yet. Tomosu scans the repository and each pull request for that second half:

These findings feed the Production Reliability Index, alongside runtime signals, so a pull request that adds a module-level Map keyed by user ID is flagged before merge. Whether it matters is then a question for runtime data, not a guess.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

What does OOMKilled mean in Kubernetes?

OOMKilled means the Linux kernel killed a container because the memory charged to its cgroup reached the limit set by resources.limits.memory and could not be reclaimed. The kernel sends SIGKILL, so the container exits with code 137 and the kubelet records the reason OOMKilled. It says where the failure happened, not which code used the memory.

Is exit code 137 always an OOM kill?

No. Exit code 137 means the process received SIGKILL. That also happens when a container ignores SIGTERM and is killed after its termination grace period, for example during a rolling update or after a failed liveness probe. Only a termination reason of OOMKilled, or a kernel log line such as Memory cgroup out of memory, confirms an OOM kill.

Why is my pod OOMKilled but memory looks low?

Usually because the dashboard shows the heap instead of the cgroup working set, a short spike fell between metric scrapes, the graph averages several containers or replicas, or a sidecar was the container that was killed. Files written to an emptyDir volume with medium Memory also count against the limit without appearing in the process RSS.

What is the difference between OOMKilled and Evicted?

OOMKilled is the kernel killing one container that reached its own memory limit. Evicted is the kubelet removing a whole pod because the node is low on memory, and it ranks pods by how far their usage exceeds their memory requests. A pod can be evicted without ever reaching its limit, so the fix for eviction is usually accurate requests.

Does page cache count against the Kubernetes memory limit?

Yes, file cache is charged to the container cgroup, but the kernel reclaims clean cache before it kills a process, so ordinary page cache rarely causes an OOM kill by itself. tmpfs files, including emptyDir with medium Memory and /dev/shm, are different: they cannot be reclaimed without swap and count fully against the limit.

What should --max-old-space-size be in a container?

Set it explicitly and below the container memory limit, leaving room for Buffers, native memory, thread stacks, and code. The right gap depends on how much the service holds outside the heap, so measure external and arrayBuffers from process.memoryUsage() first. A cap above the limit turns a readable heap out-of-memory error into a silent OOM kill.

How do I find a Node.js memory leak in production?

Confirm the heap is the layer that grows by exporting heapUsed, external, arrayBuffers, and RSS separately. Then take two or three heap snapshots some time apart, using --heapsnapshot-signal or v8.writeHeapSnapshot(), and compare them in Chrome DevTools to find the type whose retained count grows. Follow its retainers back to the module-level structure that holds it.

Should I just raise the memory limit?

Only when the evidence shows a stable plateau near the limit: working set does not ramp with uptime, kills do not correlate with one endpoint or payload, and the heap cap is consistent with the new limit. If memory grows without a bound, a higher limit only makes the restarts less frequent and the leak harder to notice.


An OOM kill is the last step of a longer story: a structure that kept growing, or a request that held everything at once. Tomosu maps those paths in the code before the limit finds them. Assess your repository →