A Kubernetes OOMKilled status tells you the kernel killed a container because it reached its memory limit. It does not tell you why the container needed that memory. Before you raise the limit, you need to know whether the process is leaking, spiking on specific work, using memory outside the heap your dashboard shows, or simply running with a limit that was never big enough.
OOMKilled (exit code 137) means the Linux kernel sent SIGKILL to a container whose cgroup memory usage reached its resources.limits.memory and could not be reclaimed. The limit is where the failure happened, not necessarily the cause. Common causes:
- Leak or unbounded growth: memory climbs with uptime and a restart resets it.
- Load-driven spike: one request, job, or payload allocates far more than normal.
- Memory outside the heap: Buffers, native allocations, thread stacks, or tmpfs files.
- Limit too low: stable, legitimate usage sits too close to the limit.
- Not an OOM kill at all: node-pressure eviction or a runtime heap limit.
The rest of this guide shows how to tell those apart with evidence: the container’s termination state, cgroup counters, the right memory metric, heap versus RSS, and a map of the code paths that accumulate memory. Examples use Node.js, with JVM equivalents where they differ.
What does OOMKilled (exit code 137) mean in Kubernetes?
Every container with a memory limit runs inside a Linux cgroup whose ceiling is set from resources.limits.memory (memory.max on cgroup v2, memory.limit_in_bytes on v1). When the processes in that cgroup try to allocate past the ceiling, the kernel first tries to reclaim memory, mostly clean file cache. If it cannot free enough, the cgroup OOM killer picks a process in that cgroup and sends it SIGKILL. The kubelet notices the container died from an OOM kill and records it:
State: Running
Started: Tue, 29 Sep 2026 10:42:11 +0000
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
Started: Tue, 29 Sep 2026 06:15:02 +0000
Finished: Tue, 29 Sep 2026 10:42:09 +0000
Restart Count: 14
Limits:
memory: 512Mi
Requests:
memory: 256Mi
Exit code 137 is 128 + 9: the process was terminated by signal 9, SIGKILL. Because SIGKILL cannot be caught, the application gets no chance to log anything. That is why an OOM kill usually leaves the application logs looking normal right up to the last line.
Any SIGKILL produces exit code 137. A container that ignores SIGTERM during a failed liveness probe or a rolling update is killed after terminationGracePeriodSeconds and also exits 137, but its reason is Error, not OOMKilled. Trust the reason field, not the exit code alone.
On the node, the kernel logs every OOM kill. The constraint field tells you whether a container limit (CONSTRAINT_MEMCG) or the whole node ran out (CONSTRAINT_NONE):
oom-kill:constraint=CONSTRAINT_MEMCG,nodemask=(null),cpuset=...,oom_memcg=/kubepods/burstable/pod...,task=node,pid=48213,uid=1000
Memory cgroup out of memory: Killed process 48213 (node) total-vm:1893420kB,
anon-rss:515204kB, file-rss:31120kB, shmem-rss:0kB, UID:1000 pgtables:1680kB oom_score_adj:984
The anon-rss and shmem-rss values are the first clue to what filled the limit. Here almost all of the 512 Mi was anonymous memory: heap, Buffers, or native allocations, not files.
OOMKilled, Evicted, or heap out of memory: which one happened?
Three different mechanisms end a container because of memory, and teams regularly confuse them. They are enforced by different components, they leave different evidence, and they need different fixes.
Node-pressure eviction
When the node as a whole runs low, the kubelet evicts pods before the kernel has to step in. It watches memory.available (node capacity minus the node’s working set) against eviction thresholds; the default hard threshold on Linux is memory.available<100Mi. Evicted pods end in phase Failed with reason Evicted, and the kubelet ranks candidates by whether their usage exceeds their requests, then by priority, then by how far above requests they are. A pod can be evicted without ever reaching its own limit. The fix is usually honest memory requests, not a higher limit. See the Kubernetes docs on node-pressure eviction.
If memory runs out faster than the kubelet can evict, the kernel’s system-wide OOM killer acts. It picks victims using oom_score_adj, which the kubelet sets by QoS class: -997 for Guaranteed, 1000 for BestEffort, and a value in between for Burstable that depends on the memory request. That kill is also reported as OOMKilled, but the kernel log shows CONSTRAINT_NONE and the node usually shows memory pressure at the same time.
Runtime heap limit
Node.js and the JVM enforce their own heap caps inside the process. When V8 cannot grow the heap past --max-old-space-size, Node prints FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory and aborts. The container exits with 134 and reason Error. This is the better failure: you get a message, a stack, and optionally a heap snapshot. It also means the heap cap, not the container limit, was the binding constraint.
# Termination reason and exit code for every container, including sidecars
kubectl get pod api-7c9f8d6b5-x2k4q -o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.lastState.terminated.reason}{"\t"}{.lastState.terminated.exitCode}{"\n"}{end}'
# Evictions and node memory pressure
kubectl get events -A --field-selector reason=Evicted
kubectl describe node <node> | grep -A2 MemoryPressure
# Inside a running container on cgroup v2: limit, current usage, peak, and OOM counters
cat /sys/fs/cgroup/memory.max /sys/fs/cgroup/memory.current /sys/fs/cgroup/memory.peak
cat /sys/fs/cgroup/memory.events # high, max, oom, oom_kill counts
grep -E '^(anon|file|shmem|kernel|sock|active_file|inactive_file) ' /sys/fs/cgroup/memory.stat
On cgroup v1 the equivalents are memory.limit_in_bytes, memory.usage_in_bytes, memory.max_usage_in_bytes, and memory.failcnt. memory.peak needs a recent kernel (5.19 or later).
What counts against a Kubernetes container memory limit?
Everything the container’s processes cause the kernel to charge to its cgroup counts. That is more than the heap your application dashboard shows, and more than most people expect:
- Heap: the V8 or JVM heap, including space the runtime reserved but is not using.
- Off-heap and native memory: Node.js
BufferandArrayBufferbacking stores, JVM direct buffers and metaspace, native addons, compression and TLS libraries, allocator fragmentation. - Thread stacks and code: worker threads, JIT-compiled code, and runtime bookkeeping.
- tmpfs: files written to an
emptyDirwithmedium: Memory, to/dev/shm, or to any path backed by tmpfs. The Kubernetes volumes documentation states that files you write there count against the container’s memory limit. - Page cache and kernel memory: file cache from reads and writes, plus kernel structures such as socket buffers. Clean cache is reclaimable; the kernel drops it before it kills anything.
Two metric details matter during an investigation. First, container_memory_working_set_bytes is usage minus inactive_file. It is what kubectl top and the eviction manager use, and it is the best single proxy for “how close is this container to being killed”. It still includes active file cache, which the kernel can reclaim, so a working set near the limit does not guarantee a kill is coming. Second, container_memory_rss counts anonymous memory only. It excludes tmpfs, which is accounted as shared memory (shmem in memory.stat). If working set is high and RSS is not, look at tmpfs and file cache.
Why is my pod OOMKilled but memory looks low?
This is the most common question in Stack Overflow threads about OOMKilled, and there are five usual explanations. Check them in this order.
- You are looking at the wrong metric. A heap graph,
process.memoryUsage().heapUsed, or JVM “heap used” is a fraction of what the cgroup charges. Compare working set to the limit, per container. - The spike fell between scrapes. Prometheus typically scrapes every 15 to 60 seconds. A request that allocates 400 MB and gets killed within two seconds never appears in a stored sample.
- You are looking at the wrong container. Pod-level graphs sum or average containers. The killed container may be a service mesh sidecar, a log shipper, or an init container with its own small limit.
- Aggregation hid it.
avgacross replicas, oravg_over_timeover five minutes, flattens the one pod that hit the limit. Usemax. - The memory is not in the process. tmpfs files in an
emptyDirwithmedium: Memoryare charged to the container even though no process shows them in its RSS.
For spikes, rely on counters rather than gauges. The restart count and last termination reason from kube-state-metrics are exact, and memory.peak or memory.events inside the container record the high-water mark and every OOM event regardless of scrape timing. Useful queries:
# Working set as a fraction of the limit, per container (max, not avg)
max by (namespace, pod, container) (container_memory_working_set_bytes{container!="", container!="POD"})
/ on (namespace, pod, container)
max by (namespace, pod, container) (kube_pod_container_resource_limits{resource="memory"})
# Containers whose last termination was an OOM kill and that restarted recently
kube_pod_container_status_last_terminated_reason{reason="OOMKilled"} == 1
and on (namespace, pod, container)
increase(kube_pod_container_status_restarts_total[1h]) > 0
# Heap vs working set gap (needs an app-level heap metric, e.g. prom-client defaults)
max by (pod) (container_memory_working_set_bytes{container="api"})
- on (pod) max by (pod) (nodejs_heap_size_used_bytes)
Kubernetes OOMKilled causes: memory leak, load spike, or limit too low?
Once you have confirmed a genuine OOM kill in your own container, the question is what shape the memory had before it died. The four shapes below account for most incidents, and each points at a different fix.
1. Leak or unbounded growth
Memory rises with uptime, independent of traffic, and never returns to its starting level. Restart intervals are roughly regular, and a deploy “fixes” it for a while. Strictly, garbage-collected languages rarely leak in the C sense. What they have is unbounded retention: a structure that is still reachable and keeps growing, such as a module-level cache, a listener list, or a map of sessions nobody removes.
2. Load-driven spike
Baseline is healthy, and kills line up with a particular endpoint, batch job, tenant, or input size. The code is usually buffering a whole payload: reading a full query result, building a whole CSV or JSON document in memory, decompressing an upload, or running unbounded Promise.all over a large list. Memory per request × concurrent requests exceeds the limit.
3. Limit too low for the steady state
Working set climbs during warm-up to a plateau close to the limit and stays there. Nothing grows over time; there is just no headroom. This is the only cause where raising the limit is the fix, and it is also common when the heap cap was set larger than the container, so the runtime happily fills memory the cgroup does not have.
4. Memory outside the heap
The heap graph is flat or small, while RSS or working set climbs. Suspects: Buffers retained by a stream that is never consumed, a native module, allocator fragmentation, unclosed compression streams, JVM direct buffers or thread growth, and files written to tmpfs.
The decision tree below puts the checks in the order that eliminates the most confusion first.
| Signal | Leak | Load spike | Outside the heap | Limit too low |
|---|---|---|---|---|
| Working set over days | Ramps with uptime, sawtooth at each restart | Flat, with sharp jumps | Ramps or steps | Plateau near the limit |
| After traffic drops | Stays high | Returns to baseline | Often stays high | Stays at plateau |
| Heap vs RSS | Heap grows with RSS | Heap spikes if JS objects; RSS if Buffers | Heap flat, RSS or shmem grows | Both stable |
| Correlates with | Uptime, distinct keys or users seen | An endpoint, job, tenant, or input size | Streams, uploads, native modules, tmpfs writes | Every pod, from startup |
| Heap snapshots | Retained count of one type grows between snapshots | Large transient arrays or strings | Little change | Little change |
| Kernel log | anon-rss near limit | anon-rss near limit | shmem-rss or anon outside heap | anon-rss near limit |
How do you find a Node.js memory leak in production?
A Node.js memory leak in production is hard mostly because the evidence lives in three layers: V8’s heap, the process’s memory outside the heap, and the cgroup. Start by making the layers visible, then take heap snapshots only when you know the heap is the layer growing.
Step 1: Set the heap cap below the container limit
--max-old-space-size sets the V8 old-generation heap limit in megabytes. The default is derived from the memory Node detects, and depending on version it may not match your container limit. Set it explicitly and leave room for everything that is not heap: Buffers, native memory, stacks, and code.
containers:
- name: api
image: registry.example.com/api:1.42.0
env:
- name: NODE_OPTIONS
value: "--max-old-space-size=640 --heapsnapshot-signal=SIGUSR2"
resources:
requests:
memory: "768Mi"
limits:
memory: "1Gi" # ~384 MiB left for Buffers, native, stacks, code
With the cap below the limit, a heap leak ends as a JavaScript heap out-of-memory crash with a message and a stack. With the cap above the limit, the same leak ends as a silent OOM kill. How much headroom to leave depends on how much the service holds outside the heap, which is the next thing to measure.
Step 2: Export heap and non-heap memory separately
const MB = 1024 * 1024;
setInterval(() => {
const m = process.memoryUsage();
logger.info({
rss_mb: Math.round(m.rss / MB), // resident set: what the kernel sees
heap_total_mb: Math.round(m.heapTotal / MB), // V8 heap reserved
heap_used_mb: Math.round(m.heapUsed / MB), // live JS objects + garbage not yet collected
external_mb: Math.round(m.external / MB), // C++ objects bound to JS, incl. Buffers
array_buffers_mb: Math.round(m.arrayBuffers / MB), // Buffer / ArrayBuffer backing stores
}, 'memory');
}, 15_000).unref();
If you use prom-client, its default metrics already export heap and RSS figures. Read them together: heapUsed growing points at retained JavaScript objects. external or arrayBuffers growing points at Buffers. RSS growing while all of those are flat points at native memory or allocator fragmentation.
Step 3: Compare heap snapshots
A single heap snapshot shows what is big. Two or three snapshots taken minutes or hours apart show what is growing, which is what you need for a leak. With --heapsnapshot-signal=SIGUSR2 set, send the signal to the Node process to write a .heapsnapshot file to the working directory. Alternatively call v8.writeHeapSnapshot() from an admin endpoint, or use --heapsnapshot-near-heap-limit=N to capture automatically as the heap approaches its cap.
# Take the pod out of rotation first if you can: the process pauses while writing
kubectl exec api-7c9f8d6b5-x2k4q -- sh -c 'kill -USR2 $(pgrep -o node)'
kubectl exec api-7c9f8d6b5-x2k4q -- ls -la /app/*.heapsnapshot
kubectl cp api-7c9f8d6b5-x2k4q:/app/Heap.20260929.104512.1.0.001.heapsnapshot ./snap-1.heapsnapshot
# Repeat later, then load both in Chrome DevTools > Memory and use the Comparison view
The Node.js v8 documentation notes that creating a heap snapshot requires memory about twice the size of the heap at the time. A process using 600 MB of heap in a 1 GiB container may be killed mid-snapshot. Capture on a pod with a temporarily raised limit, or capture early in the growth curve and compare, and make sure the snapshot is written to a disk-backed volume, not tmpfs.
In the Comparison view, sort by Size Delta or # Delta and open the retainers of the growing type. The retainer chain usually ends at a module-level variable: a Map, an array, an emitter’s listener list, or a closure captured by a timer. That variable is the accumulation path.
If snapshots show no heap growth while RSS climbs, stop looking at JavaScript objects. Check arrayBuffers for Buffers retained by streams, native addons, and whether the service writes to a memory-backed volume.
The same questions for the JVM
The JVM has the same layers with different names, and the same trap: the heap is only part of what the cgroup charges.
| Question | Node.js | JVM |
|---|---|---|
| Heap cap | --max-old-space-size | -Xmx, or -XX:MaxRAMPercentage (default 25% of the container limit on container-aware JDKs) |
| Heap exhausted | “JavaScript heap out of memory”, process aborts | OutOfMemoryError: Java heap space; add -XX:+ExitOnOutOfMemoryError so the pod restarts cleanly |
| Outside the heap | Buffers, native addons, stacks | Metaspace, thread stacks, code cache, direct buffers (-XX:MaxDirectMemorySize), GC structures |
| Non-heap breakdown | process.memoryUsage() | -XX:NativeMemoryTracking=summary, then jcmd <pid> VM.native_memory summary |
| Heap dump | Heap snapshot, DevTools Comparison | jcmd <pid> GC.heap_dump, -XX:+HeapDumpOnOutOfMemoryError |
HeapDumpOnOutOfMemoryError fires only when the JVM itself throws OutOfMemoryError. A cgroup OOM kill is SIGKILL, so the JVM never gets to write anything. See the Oracle Native Memory Tracking guide.
For Go services the equivalent lever is GOMEMLIMIT, a soft limit that makes the garbage collector work harder as the process approaches it. Set it below the container limit for the same reason.
Find the code paths that accumulate memory
Runtime evidence tells you that memory grows and, with snapshots, which type grows in one process. It does not give you the list of code paths that can grow without a bound, including the ones that have not caused an incident yet. That list comes from reading the code for a specific question: for every structure that lives longer than a request, what limits its size?
| Accumulation path | What grows | Bounded by | Fix |
|---|---|---|---|
Module-level Map or object used as a cache | One entry per distinct key | Nothing: distinct users, IDs, or URLs | LRU with max and TTL, or an external cache |
| Listener added per request or connection | Listener closures and everything they capture | Nothing until removed | Remove on close, use once, or an AbortSignal |
| Whole-payload buffering | Full result sets, documents, uploads | Input size × concurrency | Stream with backpressure; cap body sizes |
| Unbounded concurrency | In-flight work for every item | Length of the input list | Concurrency limit or batching |
| In-memory queue or retry buffer | Backlog while a dependency is slow | Downstream latency | Bounded queue, shed or reject when full |
| Metrics labels with user or request IDs | Time series in the client registry | Distinct label values | Bounded label sets only |
Files on emptyDir medium: Memory | tmpfs pages charged to the container | sizeLimit, if set | Disk-backed volume, cleanup, sizeLimit |
Unbounded cache keyed by request data
const cache = new Map();
export async function getProfile(id) {
if (cache.has(id)) return cache.get(id);
const profile = await db.profiles.findById(id);
cache.set(id, profile); // never evicted: size = number of distinct users since start
return profile;
}
import { LRUCache } from 'lru-cache';
const cache = new LRUCache({ max: 10_000, ttl: 5 * 60_000 }); // at most 10k entries, 5 min each
export async function getProfile(id) {
const hit = cache.get(id);
if (hit) return hit;
const profile = await db.profiles.findById(id);
cache.set(id, profile);
return profile;
}
This is the classic leak shape: in a load test with 50 test users, the map stops at 50 entries and looks fine. In production it grows with the number of distinct users since the last restart.
Listener registered per request
// Before: every SSE client adds a listener that is never removed.
app.get('/events', (req, res) => {
res.setHeader('Content-Type', 'text/event-stream');
const onUpdate = (msg) => res.write(`data: ${JSON.stringify(msg)}\n\n`);
bus.on('update', onUpdate); // retains res and its buffers after the client leaves
});
// After: remove the listener when the connection closes.
app.get('/events', (req, res) => {
res.setHeader('Content-Type', 'text/event-stream');
const onUpdate = (msg) => res.write(`data: ${JSON.stringify(msg)}\n\n`);
bus.on('update', onUpdate);
req.on('close', () => bus.off('update', onUpdate)); // bounded by open connections
});
Node warns about this pattern with MaxListenersExceededWarning: Possible EventEmitter memory leak detected once an emitter has more than 10 listeners for one event. Teams often silence it with setMaxListeners(0). Treat that call in a diff as a question, not a fix.
Buffering a whole payload
app.get('/export', async (req, res) => {
const { rows } = await pool.query(EXPORT_SQL, [req.query.since]); // entire result set
const csv = rows.map(toCsvLine).join('\n'); // a second full copy
res.type('text/csv').send(csv);
});
import QueryStream from 'pg-query-stream';
import { pipeline } from 'node:stream/promises';
app.get('/export', async (req, res) => {
const client = await pool.connect();
try {
const rows = client.query(new QueryStream(EXPORT_SQL, [req.query.since], { batchSize: 500 }));
res.type('text/csv');
await pipeline(rows, toCsv(), res); // toCsv(): object-mode Transform; backpressure end to end
} finally {
client.release();
}
});
The first version works in staging, where the export returns 2,000 rows. It becomes a load-driven spike the first time a large tenant exports two years of data, and it multiplies when two such exports run at once. Upload handlers have the same shape: a body parser configured with a very large limit, or await file.arrayBuffer() on an unbounded upload.
Fixes, matched to the cause
| Cause | Fix | What not to do |
|---|---|---|
| Leak or unbounded growth | Find the retaining structure from snapshot comparison, then bound it: size limits, TTLs, listener removal. Find the sibling paths written the same way. | Raise the limit. A leak fills any limit, only later. |
| Load-driven spike | Stream instead of buffer, cap request and upload sizes, limit concurrency per pod. | Size every pod for the largest tenant’s worst request. |
| Memory outside the heap | Measure the layer (arrayBuffers, NMT, shmem), then fix streams, native usage, or tmpfs writes. | Tune the heap flag. It does not govern this memory. |
| Runtime heap limit | Treat as a leak or spike investigation with a free stack trace. Raise the cap only if the heap plateau is legitimate and the container has room. | Set the heap cap above the container limit. |
| Node-pressure eviction | Set memory requests from measured usage; use Guaranteed QoS for critical pods; check node allocatable. | Raise limits without raising requests. |
| Limit too low | Raise the limit from the measured peak (memory.peak, max working set) plus headroom, and keep the heap cap consistent with it. | Remove the limit. The node then absorbs the problem. |
A bigger memory limit turns a leak that restarts every four hours into one that restarts every eight.
Evidence to require before any fix
Memory limit bumps are the easiest change to approve in a review: one line of YAML, obviously safe, and the alerts stop for a while. Whether the proposal comes from a teammate, an answer online, or an AI coding assistant, ask for the evidence that matches it. This is the same discipline as the connection pool timeouts in part one of this series: the error names the resource that ran out, not the code that used it up.
- Termination reason is OOMKilled for the app container
- Working set plateaus; it does not ramp with uptime
- No correlation with one endpoint or input size
- Heap cap is consistent with the new limit
- Growth continues at low traffic
- Snapshot comparison shows a growing retained type
- The retaining structure is named in code
- What bounds it, or why nothing does, is shown
- Kernel log shows
CONSTRAINT_NONEor eviction events - Node
MemoryPressureat the time - Requests compared with real usage on that node
For how this fits into release decisions more broadly, see Pre Merge Reliability Analysis and Production Reliability vs Observability. Observability tells you a pod died; reliability analysis asks which code made it likely.
How Tomosu helps
Runtime data and repository review answer different halves of the OOMKilled question. Metrics, the kernel log, and heap snapshots confirm that memory grows and which layer grows. A review of the repository finds which code paths can grow without a bound, including those that have not caused an incident yet. Tomosu scans the repository and each pull request for that second half:
- Risky accumulation paths: long-lived maps and caches without a size or TTL bound, listeners and timers registered per request without removal, and in-memory queues that grow while a dependency is slow.
- Whole-payload buffering: handlers that load full result sets, build whole documents in memory, or accept unbounded uploads, and unbounded concurrency over input lists.
- Configuration that hides the failure: heap caps above the container limit, memory-backed volumes without
sizeLimit, and limit changes that arrive without a matching change in code. - Blast radius: which endpoints and jobs reach each path, so an unbounded cache on the checkout path is weighed differently from one in an admin tool.
- The runtime evidence to collect: which metric, snapshot comparison, or cgroup counter would confirm whether a flagged path actually matters in production.
These findings feed the Production Reliability Index, alongside runtime signals, so a pull request that adds a module-level Map keyed by user ID is flagged before merge. Whether it matters is then a question for runtime data, not a guess.
Scan your repository with Tomosu →
Key takeaways
- OOMKilled means the kernel killed a container at its cgroup memory limit. Exit code 137 alone only means
SIGKILL; check the reason field. - OOMKilled, Evicted, and a runtime “heap out of memory” are three different mechanisms with three different fixes.
- The limit counts heap, off-heap Buffers, native memory, stacks, tmpfs, and file cache. Compare working set, not heap, to the limit.
- “Memory looks low” usually means the wrong metric, the wrong container, averaging, or a spike between scrapes.
- The shape over days separates a leak, a spike, off-heap growth, and a limit that is simply too low.
- Set the heap cap below the container limit so heap exhaustion fails with a stack instead of a silent kill.
- Snapshots find one growing structure; a code review for unbounded accumulation finds the rest.
Frequently asked questions
What does OOMKilled mean in Kubernetes?
OOMKilled means the Linux kernel killed a container because the memory charged to its cgroup reached the limit set by resources.limits.memory and could not be reclaimed. The kernel sends SIGKILL, so the container exits with code 137 and the kubelet records the reason OOMKilled. It says where the failure happened, not which code used the memory.
Is exit code 137 always an OOM kill?
No. Exit code 137 means the process received SIGKILL. That also happens when a container ignores SIGTERM and is killed after its termination grace period, for example during a rolling update or after a failed liveness probe. Only a termination reason of OOMKilled, or a kernel log line such as Memory cgroup out of memory, confirms an OOM kill.
Why is my pod OOMKilled but memory looks low?
Usually because the dashboard shows the heap instead of the cgroup working set, a short spike fell between metric scrapes, the graph averages several containers or replicas, or a sidecar was the container that was killed. Files written to an emptyDir volume with medium Memory also count against the limit without appearing in the process RSS.
What is the difference between OOMKilled and Evicted?
OOMKilled is the kernel killing one container that reached its own memory limit. Evicted is the kubelet removing a whole pod because the node is low on memory, and it ranks pods by how far their usage exceeds their memory requests. A pod can be evicted without ever reaching its limit, so the fix for eviction is usually accurate requests.
Does page cache count against the Kubernetes memory limit?
Yes, file cache is charged to the container cgroup, but the kernel reclaims clean cache before it kills a process, so ordinary page cache rarely causes an OOM kill by itself. tmpfs files, including emptyDir with medium Memory and /dev/shm, are different: they cannot be reclaimed without swap and count fully against the limit.
What should --max-old-space-size be in a container?
Set it explicitly and below the container memory limit, leaving room for Buffers, native memory, thread stacks, and code. The right gap depends on how much the service holds outside the heap, so measure external and arrayBuffers from process.memoryUsage() first. A cap above the limit turns a readable heap out-of-memory error into a silent OOM kill.
How do I find a Node.js memory leak in production?
Confirm the heap is the layer that grows by exporting heapUsed, external, arrayBuffers, and RSS separately. Then take two or three heap snapshots some time apart, using --heapsnapshot-signal or v8.writeHeapSnapshot(), and compare them in Chrome DevTools to find the type whose retained count grows. Follow its retainers back to the module-level structure that holds it.
Should I just raise the memory limit?
Only when the evidence shows a stable plateau near the limit: working set does not ramp with uptime, kills do not correlate with one endpoint or payload, and the heap cap is consistent with the new limit. If memory grows without a bound, a higher limit only makes the restarts less frequent and the leak harder to notice.
An OOM kill is the last step of a longer story: a structure that kept growing, or a request that held everything at once. Tomosu maps those paths in the code before the limit finds them. Assess your repository →