A cache stampede starts with a cache that is working exactly as configured. A popular key expires, every request that arrives in the next few hundred milliseconds misses at once, and each one runs the same expensive query. The hit ratio on the dashboard still looks healthy. The database does not.
A cache stampede (also called a cache miss storm, dogpile, or thundering herd) happens when a hot key expires and many concurrent requests miss together, each recomputing the same value against the database. The extra load is roughly request rate × recompute time, per key. To prevent a cache stampede, make sure only one caller recomputes a value and nobody waits on an empty cache:
- Coalesce requests: single-flight in process, plus a short Redis lock (
SET key token NX PX) across instances. - Spread expiry: add TTL jitter so keys written together do not expire together.
- Refresh before expiry: soft TTL with stale-while-revalidate, or probabilistic early expiration (XFetch).
- Avoid cold starts: warm caches before traffic, never flush a shared cache during a deploy, and cache “not found” results briefly.
None of these needs a new system. They are a few lines of cache access code each, and so are the changes that break them. That is why this post ends with the pull request patterns that turn a quiet cache into an incident.
What is a cache stampede?
Most services cache with the cache-aside pattern: read the key, and on a miss, compute the value from the database and write it back with a TTL. That code is correct for one request. It has no answer for the case where many requests miss the same key in the same instant.
def get_product(product_id):
key = f"product:{product_id}"
cached = r.get(key)
if cached is not None:
return json.loads(cached)
# every concurrent miss reaches this line at the same time
product = db.fetch_product_with_pricing(product_id) # 400 ms, 6 joins
r.set(key, json.dumps(product), ex=300)
return product
Between the moment the key expires and the moment the first recompute writes it back, the key does not exist. Every request in that window takes the miss branch. If the key is read 2,000 times a second and the query takes 400 ms, roughly 800 requests start the same query before the first one finishes. Those 800 queries return identical rows, and 799 of them are pure waste.
The pattern has several names. The 2015 VLDB paper on the subject calls it “cache stampede (also called dog-piling, cache miss storm, or cache choking)”. Ops teams often say thundering herd. Whatever the name, there are two ingredients: a hot key (high read rate) and an expensive recompute (slow query, fan-out to several services, or a heavy aggregation). A cold key with a slow query is fine. A hot key with a 2 ms primary-key lookup is usually fine. The product of the two is what hurts.
Why does a cache miss storm overload the database?
Developer questions on Stack Overflow describe the same surprise: the cache hit ratio is above 99%, yet the database periodically saturates and the slow query log fills with one statement. A cache is a load absorber. The database behind it has been sized for the traffic that gets past the cache, not for the traffic that arrives at the front door. When a hot key misses, the full front-door rate for that key lands on a backend that was provisioned for a fraction of it.
Three details make the damage larger than the arithmetic suggests:
- The feedback loop. Concurrent copies of the same query contend for the same pages, locks, and CPU. The recompute time goes up, the miss window widens, and more requests pile in. The stampede on one key can slow the database enough to push other keys into their own stampedes.
- Connection pools fill first. Hundreds of requests each holding a database connection for a slow query is exactly the starvation pattern in HikariCP: Connection Is Not Available, Request Timed Out. Unrelated endpoints that share the pool start timing out.
- Retries multiply it. Requests that time out during the stampede get retried by clients or gateways, which adds more misses into the same window. That is the mechanism covered in How to Spot a Retry Storm Before Merge.
A stampede only happens when expiry lines up with a traffic peak. At 3 a.m. the same key expires harmlessly. That is why these incidents look random, recur on a schedule nobody has written down, and survive load tests that do not hold traffic steady across a TTL boundary.
Synchronized expiry: what causes a cache miss storm across many keys?
One hot key is a spike. Thousands of keys expiring in the same second is an outage. Synchronized expiry happens whenever many keys are written at the same moment with the same TTL, because they will then expire at the same moment too:
- Warm-up jobs. A job preloads 10,000 product keys at deploy time with
EX 3600. One hour later they all expire together, often at the same point in the next traffic cycle. - Batch refreshes. A cron job rebuilds a leaderboard or catalog every hour and writes every entry with the same TTL.
- Restarts and deploys. In-process caches (a local LRU, Caffeine, a module-level dict) are empty in every new pod. A rolling deploy of 40 pods is 40 cold caches, each with its own burst of misses.
- Flushes and key-version bumps.
FLUSHALLin a migration, a new Redis cluster, or changing a key prefix fromv1:tov2:all make every key miss at once. A prefix change is a flush that does not look like one in review.
import random
def jittered_ttl(base_seconds: int, spread: float = 0.10) -> int:
# 3600 with spread 0.10 -> uniform in [3240, 3960]
delta = int(base_seconds * spread)
return base_seconds + random.randint(-delta, delta)
r.set(key, payload, ex=jittered_ttl(3600))
Jitter is cheap and belongs in the shared cache helper, not at each call site, so that nobody can forget it. It does not protect a single hot key: that key still expires at some instant and still stampedes. For that you need coalescing.
How do you prevent a cache stampede with request coalescing?
Request coalescing, usually called single-flight, means that for any given key only one caller runs the recompute. Everyone else who misses the same key while it is running waits for that result instead of starting their own query. It turns N identical queries into one.
In Go, golang.org/x/sync/singleflight does exactly this. Group.Do(key, fn) runs fn once per key at a time and hands the same result to every concurrent caller; the third return value, shared, reports whether the result went to more than one caller.
var sf singleflight.Group
func (s *Store) GetProduct(ctx context.Context, id string) (*Product, error) {
key := "product:" + id
if b, err := s.rdb.Get(ctx, key).Bytes(); err == nil {
return decode(b)
} else if err != redis.Nil {
return nil, err // Redis down is a different problem; do not stampede the DB
}
// Only one goroutine per key in this process runs the loader.
ch := sf.DoChan(key, func() (any, error) {
// Detach from the leader's request context so one cancelled
// client does not fail every follower; bound it instead.
lctx, cancel := context.WithTimeout(context.WithoutCancel(ctx), 2*time.Second)
defer cancel()
p, err := s.db.FetchProductWithPricing(lctx, id)
if err != nil {
return nil, err
}
_ = s.rdb.Set(lctx, key, encode(p), jitteredTTL(5*time.Minute)).Err()
return p, nil
})
select {
case res := <-ch:
if res.Err != nil {
return nil, res.Err
}
return res.Val.(*Product), nil
case <-ctx.Done():
return nil, ctx.Err() // this caller gives up; the load continues for others
}
}
Three details in that sketch matter in production. DoChan plus select lets each caller respect its own deadline without abandoning the shared load. context.WithoutCancel (Go 1.21+) stops the leader’s client disconnect from cancelling the query that every follower is waiting on. And an error is shared too: if the loader fails, every follower gets the same error, which is usually what you want during an outage, because it stops N callers from retrying N queries.
Other stacks have the same primitive under different names. In the JVM, Caffeine’s LoadingCache and Guava’s LoadingCache coalesce concurrent loads of the same key. In Node.js, keep a Map of in-flight promises keyed by cache key and delete the entry when the promise settles. At the HTTP layer, NGINX has proxy_cache_lock, which lets only one request populate a cache element at a time.
Single-flight deduplicates within one process. With 40 application instances, a hot key miss still produces up to 40 identical queries, one per instance. That is usually a 20× to 100× improvement and often enough. When it is not, add a cross-instance guard in Redis.
How do you prevent a Redis cache stampede with SET NX PX?
To coalesce across instances, the instance that wins the right to recompute takes a short-lived lock in Redis. The standard single-instance pattern from the Redis distributed locks documentation uses SET with NX (only set if the key does not exist) and PX (expire in milliseconds), and a random token so only the owner can release it:
# Acquire: returns OK for exactly one caller, nil for everyone else
SET lock:product:42 "b7e3c1f0-instance-17" NX PX 5000
# Release only if we still own it (the token matches)
EVAL "if redis.call('get', KEYS[1]) == ARGV[1] then
return redis.call('del', KEYS[1])
else
return 0
end" 1 lock:product:42 "b7e3c1f0-instance-17"
The Redis docs note that Redis 8.4 adds DELEX key IFEQ value for the same compare-and-delete; on older versions use the script. A plain DEL is unsafe, because a slow owner whose lock already expired would delete the next owner’s lock.
The lock decides who recomputes. The harder design question is what everyone else does:
| Loser strategy | Behavior | Cost | Use when |
|---|---|---|---|
| Serve stale | Return the previous value while the winner refreshes | Readers see data up to one recompute old | You keep a stale copy (soft TTL). The best default. |
| Wait and re-read | Poll the cache key with short, jittered sleeps up to a deadline | Added latency; polling load on Redis | No stale copy exists and freshness matters |
| Fail fast or fallback | Return a default, a degraded response, or an error | Visible degradation during the window | The data is optional for the response |
| Recompute anyway | Ignore the lock after a timeout | Stampede returns if the winner is slow | Only as a last-resort bound on waiting |
Pick PX from measured recompute time. If the lock expires before the winner writes the value, a second instance acquires it and you get two recomputes. If PX is far too long and the winner crashes, nobody refreshes until it expires.
It is an efficiency lock, not a correctness lock. With a replicated Redis, the Redis docs point out that a failover can lose a freshly written lock key because replication is asynchronous, so two holders are possible. For stampede protection that only means one duplicate query, which is fine. Do not reuse the same lock to protect a write that must happen exactly once.
Losers must not hot-loop. A tight GET retry loop across hundreds of requests moves the stampede from the database to Redis. Sleep with jitter and cap the total wait.
Serve stale, refresh early: soft TTLs and XFetch
Coalescing shrinks the stampede. Refreshing before the key disappears removes the miss window altogether, because there is always a value to serve. There are two common ways to do it.
Soft TTL with stale-while-revalidate
Store the value with a logical “fresh until” timestamp that is shorter than the Redis TTL. Readers that see a fresh value return it. Readers that see a stale value still return it immediately, and one of them, chosen by the SET NX PX lock, refreshes in the background. The key only truly expires (hard TTL) if nobody reads it for a long time. This is the same idea as the HTTP stale-while-revalidate extension in RFC 5861, and NGINX’s proxy_cache_use_stale updating; Caffeine offers it in-process as refreshAfterWrite.
SOFT_TTL = 60 # seconds a value counts as fresh
HARD_TTL = 600 # seconds Redis keeps it at all
def get_product(product_id):
key = f"product:{product_id}"
raw = r.get(key)
if raw is not None:
entry = json.loads(raw)
if time.time() >= entry["fresh_until"]:
# stale: serve it now, and let exactly one caller refresh
token = uuid.uuid4().hex
if r.set(f"lock:{key}", token, nx=True, px=5000):
executor.submit(refresh, key, product_id, token)
return entry["value"]
# true miss (cold key): coalesce, see the Go example above
return load_with_single_flight(key, product_id)
def refresh(key, product_id, token):
try:
value = db.fetch_product_with_pricing(product_id)
entry = {"value": value, "fresh_until": time.time() + SOFT_TTL}
r.set(key, json.dumps(entry), ex=jittered_ttl(HARD_TTL))
finally:
release_lock(f"lock:{key}", token) # compare-and-delete script
Probabilistic early expiration (XFetch)
If you cannot run background refreshes, let readers volunteer to recompute early, with a probability that rises as expiry approaches. Vattani, Chierichetti, and Lowenstein describe this in Optimal Probabilistic Cache Stampede Prevention (PVLDB, Vol. 8, No. 8, 2015). Their XFetch algorithm stores, next to the value, the time the last recompute took (Δ) and the expiry time, and recomputes when now − Δ · β · ln(rand()) ≥ expiry, with β = 1 by default. Because ln(rand()) is negative, the left side is now plus a random, exponentially distributed head start, scaled by how long recomputing takes.
import math, random, time
def should_recompute(delta: float, expiry: float, beta: float = 1.0) -> bool:
# 1 - random.random() is in (0, 1], so log() never sees 0
return time.time() - delta * beta * math.log(1.0 - random.random()) >= expiry
Slow-to-compute values start refreshing earlier; cheap ones barely at all. A hot key is almost always refreshed by one reader shortly before it expires, and the other readers keep getting the cached value. β > 1 favors earlier recomputation. XFetch needs no lock, which makes it attractive for memcached-style caches, but it is probabilistic: under very high concurrency, more than one reader can win the draw, so it pairs well with single-flight.
Cold starts, cache warming, and negative caching
Every technique above assumes there is something in the cache to protect. Two situations break that assumption.
Cold starts after a deploy or flush
- Warm before traffic, not with traffic. Preload the known hot set (from yesterday’s access logs or a top-N list) before the instance passes its readiness check, and write each key with a jittered TTL. For in-process caches, gate the Kubernetes readiness probe on warm-up completion.
- Do not flush shared caches in a deploy. Invalidate what the change actually affects. If a schema change requires new keys, write both versions during a transition, or read-through from the old key while the new one fills.
- Treat a key prefix change as a flush. Roll it out gradually (by tenant, by percentage of traffic) and watch database load while it rolls.
- Roll slowly with in-process caches. A deployment strategy that replaces half the fleet at once doubles the number of cold caches hitting the database at the same moment.
Negative caching for keys that do not exist
A lookup for an ID that does not exist always misses, because there is nothing to cache. If a client loops on a deleted ID, or a scraper walks random IDs, every request goes to the database. Cache the absence: store a sentinel such as __none__ with a short TTL (seconds to a minute), and return “not found” when you read it. Keep the TTL short and delete the sentinel when the record is created, or users will not see new records until it expires.
How to tell a stampede from ordinary load
A stampede has a distinctive fingerprint. Check for it before concluding that the database needs to be bigger:
- One query fingerprint dominates
pg_stat_statementsor the slow query log - Many concurrent sessions run the identical statement with the same parameters
- Load spikes recur at a fixed interval that matches a TTL
INFO stats:keyspace_missesjumps whilekeyspace_hitsstays highexpired_keysspikes just before the database spike- Miss rate is concentrated on a few keys or one prefix
- Spikes line up with deploys, restarts, or warm-up jobs
- Connection pool waits rise on endpoints that read the hot key
- Traces show many parallel spans running the same loader
The correlation step is the same one used for any incident: line up the metric spike with the code and deploy timeline. How to Correlate Logs, Traces, and a Code Change walks through it, and Which Commit Caused the Production Incident? covers narrowing it to a change.
| Technique | Prevents | Does not prevent | Main caveat |
|---|---|---|---|
| TTL jitter | Mass synchronized expiry | A single hot key stampede | Must be in the shared helper |
| Single-flight | Duplicate loads within a process | One query per instance | Leader errors and timeouts are shared |
Redis SET NX PX lock | Duplicate loads across instances | Waiting, if there is no stale copy | PX sizing; not a correctness lock |
| Soft TTL / SWR | The miss window on hot keys | Cold keys and cold starts | Readers see bounded staleness |
| XFetch | Expiry-time stampedes, lock-free | Cold starts; rare double refresh | Needs Δ stored with the value |
| Warming | Cold-start storms after deploys | Keys outside the warmed set | Warm-up itself needs jitter |
| Negative caching | Repeated misses on absent IDs | Misses on real keys | Short TTL; clear on create |
The pull requests that cause cache stampedes
Stampedes rarely come from a new cache. They come from small changes to an existing one, reviewed as refactors. Each looks harmless in a diff and only becomes dangerous under concurrency and real traffic:
@@ def on_startup():
+ for product in db.top_products(limit=10_000):
+ r.set(f"product:{product.id}", json.dumps(product), ex=3600)
# every pod runs this at deploy time: 10,000 keys, one TTL,
# all expiring together one hour after the deploy
@@ def update_price(product_id, price):
db.update_price(product_id, price)
- refresh_async(f"product:{product_id}") # stale served during refresh
+ r.delete(f"product:{product_id}") # next read misses
# on a hot product with frequent price updates, every write
# now opens a miss window, and nothing coalesces the reload
Delete-on-write is a legitimate, often recommended invalidation strategy: it avoids races where an older value overwrites a newer one. The risk is removing it from a path that previously served stale data without adding coalescing on the read side. The same review question applies to the others:
- A new fixed TTL on a bulk write (warm-up, batch job, backfill) without jitter.
- Removing or bypassing a loader, for example replacing a
LoadingCache.getwith a manualgetIfPresentthen load, which silently drops per-key coalescing. - Raising the cost of recompute: a new join, an extra downstream call, or a larger payload in the loader of a hot key, which widens the miss window without touching cache code at all.
- Flushes and key renames in migrations or deploy scripts:
FLUSHALL,FLUSHDB, a new key prefix, or a new cache cluster endpoint. - A lock with no loser strategy:
SET NX PXadded, but callers that lose the race immediately query the database anyway.
None of these is visible as a bug in the changed lines. The risk comes from how hot the key is, how expensive the loader is, and how many instances run the code, which is why blast radius matters more than diff size. How to Review a Pull Request for Production Reliability Risks covers the general review method; caching is one of the categories where it pays off most.
How Tomosu helps
Tomosu analyzes the standing codebase and each pull request for production reliability risk, and cache expiry logic is a case where the risk depends on context the diff does not show. For cache stampedes, Tomosu flags and maps:
- Cache access paths and their loaders: where a key is read, what runs on a miss, which database queries or downstream calls that loader makes, and whether concurrent misses are coalesced (single-flight, a loading cache, or a lock) or not.
- Expiry changes under concurrency: a PR that writes a batch of keys with one fixed TTL during warm-up, replaces a stale-while-revalidate refresh with delete-on-write, drops a loader in favor of manual get-then-load, or flushes or renames keys in a deploy or migration.
- Blast radius: which endpoints read the affected keys and which database and connection pool the loader shares, so a change to a rarely read admin cache is weighed differently from one on the product page.
- The evidence to request before merge: the expected read rate of the key, the recompute time, the number of instances, and the loser strategy if a lock is involved.
These findings feed the Production Reliability Index alongside the Fragility, Drift, and Code Volatility signals, so a warm-up job with a fixed TTL is surfaced in review, not an hour after the deploy.
Scan your repository with Tomosu →
Key takeaways
- A cache stampede is a hot key plus an expensive recompute. Extra database load is roughly request rate × recompute time, and it feeds on itself as the database slows.
- A high hit ratio does not rule out a stampede. Look for one dominant query fingerprint,
expired_keysspikes, and a recurring interval that matches a TTL. - Jitter every TTL in the shared cache helper. Keys written together must not expire together.
- Coalesce misses: single-flight per process, and a Redis
SET NX PXlock across instances with a defined loser strategy. - Prefer serving stale over waiting: soft TTL with background refresh, or XFetch, keeps a value in the cache while it is recomputed.
- Warm before traffic, avoid flushes and prefix changes in deploys, and cache “not found” briefly.
- Review cache expiry changes for concurrency, not just correctness. The dangerous PRs look like cleanups.
Frequently asked questions
What is a cache stampede?
A cache stampede happens when a frequently read cache key expires or is deleted and many concurrent requests miss it at the same time. Each request then recomputes the same value, usually by running the same database query, so the database receives a burst of identical work roughly equal to the request rate multiplied by the recompute time. It is also called a cache miss storm, dogpile effect, or thundering herd.
How do you prevent a cache stampede?
Make sure only one caller recomputes a value and that other callers do not wait on an empty cache. Use single-flight request coalescing in each process, a short Redis lock with SET NX PX across instances, TTL jitter so keys written together do not expire together, and a soft TTL with stale-while-revalidate or probabilistic early expiration so the value is refreshed before it disappears. Warm caches before traffic and avoid flushing shared caches during deploys.
What is the difference between a cache stampede and a cache miss storm?
They describe the same failure. Cache stampede usually refers to many requests missing one hot key at once. Cache miss storm is often used for the wider case where many keys miss together, for example after a deploy, a cache flush, or a warm-up job that gave thousands of keys the same TTL. Both overload the backend with work the cache normally absorbs.
How do I prevent a Redis cache stampede with a lock?
On a miss, try SET lock:key token NX PX 5000. The caller that gets OK recomputes the value, writes it to the cache, and releases the lock with a compare-and-delete script that only deletes the lock if the token still matches. Callers that do not get the lock should serve a stale copy, or wait with jittered sleeps and re-read the cache, rather than querying the database. Set PX from measured recompute time.
Does TTL jitter prevent a cache stampede?
Jitter prevents mass synchronized expiry, where many keys written at the same time expire at the same time. It does not protect a single hot key, which still expires at some instant and can still stampede. Combine jitter with request coalescing or a soft TTL for hot keys.
What is probabilistic early expiration or XFetch?
XFetch, from the 2015 VLDB paper Optimal Probabilistic Cache Stampede Prevention by Vattani, Chierichetti, and Lowenstein, lets each reader decide at random to recompute a value before it expires. The chance rises as expiry approaches and is scaled by how long the last recompute took, so one reader usually refreshes a hot key early while others keep reading the cached value. It needs no lock.
Is a Redis SET NX lock safe enough for cache stampede protection?
Yes for this purpose, because the lock only limits duplicate work. If the lock expires early or is lost in a Redis failover, the worst case is a second recompute. Do not rely on the same single-instance lock for operations that must happen exactly once, such as charging a payment, where you need fencing or a stronger coordination mechanism.
Why does my database spike after every deploy even though the cache hit ratio is high?
Deploys often empty caches: in-process caches start cold in every new instance, and deploy scripts or migrations may flush Redis or change the key prefix. The first requests after the deploy all miss together. Warm the hot set before the instance receives traffic, roll out gradually, and avoid flushes or prefix changes as part of a routine deploy.
This is the last post in the Production Debugging series. Pool timeouts, OOMKills, retry storms, duplicate webhooks, and cache stampedes share one shape: code that is correct for a single request and wrong under concurrency, shipped in a change that looked small. Tomosu maps those paths before they reach production. Assess your repository →