Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Caching

Cache Stampede: Why a Healthy Cache Can Overload Your Database

Tomosu AI·15 min read·

A cache stampede starts with a cache that is working exactly as configured. A popular key expires, every request that arrives in the next few hundred milliseconds misses at once, and each one runs the same expensive query. The hit ratio on the dashboard still looks healthy. The database does not.

Quick answer

A cache stampede (also called a cache miss storm, dogpile, or thundering herd) happens when a hot key expires and many concurrent requests miss together, each recomputing the same value against the database. The extra load is roughly request rate × recompute time, per key. To prevent a cache stampede, make sure only one caller recomputes a value and nobody waits on an empty cache:

None of these needs a new system. They are a few lines of cache access code each, and so are the changes that break them. That is why this post ends with the pull request patterns that turn a quiet cache into an incident.

What is a cache stampede?

Most services cache with the cache-aside pattern: read the key, and on a miss, compute the value from the database and write it back with a TTL. That code is correct for one request. It has no answer for the case where many requests miss the same key in the same instant.

Naive cache-aside (Python, redis-py)stampede-prone
def get_product(product_id):
    key = f"product:{product_id}"
    cached = r.get(key)
    if cached is not None:
        return json.loads(cached)
    # every concurrent miss reaches this line at the same time
    product = db.fetch_product_with_pricing(product_id)   # 400 ms, 6 joins
    r.set(key, json.dumps(product), ex=300)
    return product

Between the moment the key expires and the moment the first recompute writes it back, the key does not exist. Every request in that window takes the miss branch. If the key is read 2,000 times a second and the query takes 400 ms, roughly 800 requests start the same query before the first one finishes. Those 800 queries return identical rows, and 799 of them are pure waste.

ONE HOT KEY EXPIRES: THE MISS WINDOW REQUESTS CACHE DATABASE hit: served from Redis miss: key gone, all requests recompute hit again identical queries in flight peak ≈ rate × recompute t0: TTL expires t0 + 400 ms: first SET miss window = recompute time
A cache stampede on a single key. The miss window lasts as long as the slowest part of the first recompute, and every request inside it becomes a database query.

The pattern has several names. The 2015 VLDB paper on the subject calls it “cache stampede (also called dog-piling, cache miss storm, or cache choking)”. Ops teams often say thundering herd. Whatever the name, there are two ingredients: a hot key (high read rate) and an expensive recompute (slow query, fan-out to several services, or a heavy aggregation). A cold key with a slow query is fine. A hot key with a 2 ms primary-key lookup is usually fine. The product of the two is what hurts.

Why does a cache miss storm overload the database?

Developer questions on Stack Overflow describe the same surprise: the cache hit ratio is above 99%, yet the database periodically saturates and the slow query log fills with one statement. A cache is a load absorber. The database behind it has been sized for the traffic that gets past the cache, not for the traffic that arrives at the front door. When a hot key misses, the full front-door rate for that key lands on a backend that was provisioned for a fraction of it.

AMPLIFICATION: WHY IT GETS WORSE, NOT BETTER QUERIES PER STAMPEDE, ONE KEY misses ≈ rate × recompute 200/s × 50 ms ≈ 10 2,000/s × 400 ms ≈ 800 2,000/s × 2 s (slowed) ≈ 4,000 More identical queries in flight DB saturates: CPU, locks, conn. pool Each recompute gets slower Miss window widens FEEDBACK LOOP Once the database slows, recomputes for other keys slow too, so their miss windows widen as well.
Illustrative arithmetic, not a benchmark. The dangerous part is the loop: load makes recompute slower, and slower recompute means more load.

Three details make the damage larger than the arithmetic suggests:

Why it is intermittent

A stampede only happens when expiry lines up with a traffic peak. At 3 a.m. the same key expires harmlessly. That is why these incidents look random, recur on a schedule nobody has written down, and survive load tests that do not hold traffic steady across a TTL boundary.

Synchronized expiry: what causes a cache miss storm across many keys?

One hot key is a spike. Thousands of keys expiring in the same second is an outage. Synchronized expiry happens whenever many keys are written at the same moment with the same TTL, because they will then expire at the same moment too:

SAME KEYS, SAME AVERAGE TTL, DIFFERENT EXPIRY SHAPE FIXED TTL: EX 3600 FOR ALL Warm-up writes 10,000 keys at deploy time T. All of them expire in the same second. 10,000 T T+3600 s one spike of 10,000 misses TTL 3600 S ± 10% JITTER ≈ 1,000 keys per 72 s slice Each key gets its own TTL, drawn uniformly from 3,240 to 3,960 s. Expiry spreads over 12 minutes, about 14 misses per second. T+3240 s T+3960 s a trickle the database absorbs
TTL jitter changes nothing about how long data stays cached on average. It only removes the moment when everything expires at once.
TTL jitter (Python)spreads expiry
import random

def jittered_ttl(base_seconds: int, spread: float = 0.10) -> int:
    # 3600 with spread 0.10 -> uniform in [3240, 3960]
    delta = int(base_seconds * spread)
    return base_seconds + random.randint(-delta, delta)

r.set(key, payload, ex=jittered_ttl(3600))

Jitter is cheap and belongs in the shared cache helper, not at each call site, so that nobody can forget it. It does not protect a single hot key: that key still expires at some instant and still stampedes. For that you need coalescing.

How do you prevent a cache stampede with request coalescing?

Request coalescing, usually called single-flight, means that for any given key only one caller runs the recompute. Everyone else who misses the same key while it is running waits for that result instead of starting their own query. It turns N identical queries into one.

FIVE CONCURRENT MISSES ON PRODUCT:42 WITHOUT COALESCING req 1 req 2 req 3 req 4 req 5 miss → query Database 5 × SELECT every miss runs the same query WITH SINGLE-FLIGHT leader req 2 req 3 req 4 req 5 one load per key product:42 Database 1 SELECT 4 followers share the leader’s result
Single-flight collapses concurrent misses on the same key into one recompute. It is per process, so each instance still sends one query.

In Go, golang.org/x/sync/singleflight does exactly this. Group.Do(key, fn) runs fn once per key at a time and hands the same result to every concurrent caller; the third return value, shared, reports whether the result went to more than one caller.

Single-flight cache read (Go, go-redis v9)coalesced
var sf singleflight.Group

func (s *Store) GetProduct(ctx context.Context, id string) (*Product, error) {
    key := "product:" + id
    if b, err := s.rdb.Get(ctx, key).Bytes(); err == nil {
        return decode(b)
    } else if err != redis.Nil {
        return nil, err // Redis down is a different problem; do not stampede the DB
    }

    // Only one goroutine per key in this process runs the loader.
    ch := sf.DoChan(key, func() (any, error) {
        // Detach from the leader's request context so one cancelled
        // client does not fail every follower; bound it instead.
        lctx, cancel := context.WithTimeout(context.WithoutCancel(ctx), 2*time.Second)
        defer cancel()
        p, err := s.db.FetchProductWithPricing(lctx, id)
        if err != nil {
            return nil, err
        }
        _ = s.rdb.Set(lctx, key, encode(p), jitteredTTL(5*time.Minute)).Err()
        return p, nil
    })

    select {
    case res := <-ch:
        if res.Err != nil {
            return nil, res.Err
        }
        return res.Val.(*Product), nil
    case <-ctx.Done():
        return nil, ctx.Err() // this caller gives up; the load continues for others
    }
}

Three details in that sketch matter in production. DoChan plus select lets each caller respect its own deadline without abandoning the shared load. context.WithoutCancel (Go 1.21+) stops the leader’s client disconnect from cancelling the query that every follower is waiting on. And an error is shared too: if the loader fails, every follower gets the same error, which is usually what you want during an outage, because it stops N callers from retrying N queries.

Other stacks have the same primitive under different names. In the JVM, Caffeine’s LoadingCache and Guava’s LoadingCache coalesce concurrent loads of the same key. In Node.js, keep a Map of in-flight promises keyed by cache key and delete the entry when the promise settles. At the HTTP layer, NGINX has proxy_cache_lock, which lets only one request populate a cache element at a time.

The limit of in-process single-flight

Single-flight deduplicates within one process. With 40 application instances, a hot key miss still produces up to 40 identical queries, one per instance. That is usually a 20× to 100× improvement and often enough. When it is not, add a cross-instance guard in Redis.

How do you prevent a Redis cache stampede with SET NX PX?

To coalesce across instances, the instance that wins the right to recompute takes a short-lived lock in Redis. The standard single-instance pattern from the Redis distributed locks documentation uses SET with NX (only set if the key does not exist) and PX (expire in milliseconds), and a random token so only the owner can release it:

Redis: acquire and release a recompute lockcross-instance
# Acquire: returns OK for exactly one caller, nil for everyone else
SET lock:product:42 "b7e3c1f0-instance-17" NX PX 5000

# Release only if we still own it (the token matches)
EVAL "if redis.call('get', KEYS[1]) == ARGV[1] then
        return redis.call('del', KEYS[1])
      else
        return 0
      end" 1 lock:product:42 "b7e3c1f0-instance-17"

The Redis docs note that Redis 8.4 adds DELEX key IFEQ value for the same compare-and-delete; on older versions use the script. A plain DEL is unsafe, because a slow owner whose lock already expired would delete the next owner’s lock.

The lock decides who recomputes. The harder design question is what everyone else does:

Loser strategyBehaviorCostUse when
Serve staleReturn the previous value while the winner refreshesReaders see data up to one recompute oldYou keep a stale copy (soft TTL). The best default.
Wait and re-readPoll the cache key with short, jittered sleeps up to a deadlineAdded latency; polling load on RedisNo stale copy exists and freshness matters
Fail fast or fallbackReturn a default, a degraded response, or an errorVisible degradation during the windowThe data is optional for the response
Recompute anywayIgnore the lock after a timeoutStampede returns if the winner is slowOnly as a last-resort bound on waiting
Caveats of a Redis recompute lock

Pick PX from measured recompute time. If the lock expires before the winner writes the value, a second instance acquires it and you get two recomputes. If PX is far too long and the winner crashes, nobody refreshes until it expires.

It is an efficiency lock, not a correctness lock. With a replicated Redis, the Redis docs point out that a failover can lose a freshly written lock key because replication is asynchronous, so two holders are possible. For stampede protection that only means one duplicate query, which is fine. Do not reuse the same lock to protect a write that must happen exactly once.

Losers must not hot-loop. A tight GET retry loop across hundreds of requests moves the stampede from the database to Redis. Sleep with jitter and cap the total wait.

Serve stale, refresh early: soft TTLs and XFetch

Coalescing shrinks the stampede. Refreshing before the key disappears removes the miss window altogether, because there is always a value to serve. There are two common ways to do it.

Soft TTL with stale-while-revalidate

Store the value with a logical “fresh until” timestamp that is shorter than the Redis TTL. Readers that see a fresh value return it. Readers that see a stale value still return it immediately, and one of them, chosen by the SET NX PX lock, refreshes in the background. The key only truly expires (hard TTL) if nobody reads it for a long time. This is the same idea as the HTTP stale-while-revalidate extension in RFC 5861, and NGINX’s proxy_cache_use_stale updating; Caffeine offers it in-process as refreshAfterWrite.

Soft TTL + background refresh (Python, redis-py)no miss window
SOFT_TTL = 60      # seconds a value counts as fresh
HARD_TTL = 600     # seconds Redis keeps it at all

def get_product(product_id):
    key = f"product:{product_id}"
    raw = r.get(key)
    if raw is not None:
        entry = json.loads(raw)
        if time.time() >= entry["fresh_until"]:
            # stale: serve it now, and let exactly one caller refresh
            token = uuid.uuid4().hex
            if r.set(f"lock:{key}", token, nx=True, px=5000):
                executor.submit(refresh, key, product_id, token)
        return entry["value"]
    # true miss (cold key): coalesce, see the Go example above
    return load_with_single_flight(key, product_id)

def refresh(key, product_id, token):
    try:
        value = db.fetch_product_with_pricing(product_id)
        entry = {"value": value, "fresh_until": time.time() + SOFT_TTL}
        r.set(key, json.dumps(entry), ex=jittered_ttl(HARD_TTL))
    finally:
        release_lock(f"lock:{key}", token)   # compare-and-delete script

Probabilistic early expiration (XFetch)

If you cannot run background refreshes, let readers volunteer to recompute early, with a probability that rises as expiry approaches. Vattani, Chierichetti, and Lowenstein describe this in Optimal Probabilistic Cache Stampede Prevention (PVLDB, Vol. 8, No. 8, 2015). Their XFetch algorithm stores, next to the value, the time the last recompute took (Δ) and the expiry time, and recomputes when now − Δ · β · ln(rand()) ≥ expiry, with β = 1 by default. Because ln(rand()) is negative, the left side is now plus a random, exponentially distributed head start, scaled by how long recomputing takes.

XFetch early-expiration check (Python)no coordination
import math, random, time

def should_recompute(delta: float, expiry: float, beta: float = 1.0) -> bool:
    # 1 - random.random() is in (0, 1], so log() never sees 0
    return time.time() - delta * beta * math.log(1.0 - random.random()) >= expiry

Slow-to-compute values start refreshing earlier; cheap ones barely at all. A hot key is almost always refreshed by one reader shortly before it expires, and the other readers keep getting the cached value. β > 1 favors earlier recomputation. XFetch needs no lock, which makes it attractive for memcached-style caches, but it is probabilistic: under very high concurrency, more than one reader can win the draw, so it pairs well with single-flight.

REFRESH BEFORE THE KEY DISAPPEARS SOFT TTL + STALE-WHILE-REVALIDATE fresh: serve from cache stale: serve + one background refresh hard expired written soft TTL hard TTL XFETCH: CHANCE A READER RECOMPUTES EARLY near zero for most of the lifetime rises within a few Δ written expiry Δ = last recompute time
Both approaches keep a value in the cache while it is being recomputed. Soft TTL coordinates with a lock; XFetch spreads the decision across readers with a random draw.

Cold starts, cache warming, and negative caching

Every technique above assumes there is something in the cache to protect. Two situations break that assumption.

Cold starts after a deploy or flush

Negative caching for keys that do not exist

A lookup for an ID that does not exist always misses, because there is nothing to cache. If a client loops on a deleted ID, or a scraper walks random IDs, every request goes to the database. Cache the absence: store a sentinel such as __none__ with a short TTL (seconds to a minute), and return “not found” when you read it. Keep the TTL short and delete the sentinel when the record is created, or users will not see new records until it expires.

How to tell a stampede from ordinary load

A stampede has a distinctive fingerprint. Check for it before concluding that the database needs to be bigger:

In the database
  • One query fingerprint dominates pg_stat_statements or the slow query log
  • Many concurrent sessions run the identical statement with the same parameters
  • Load spikes recur at a fixed interval that matches a TTL
In the cache
  • INFO stats: keyspace_misses jumps while keyspace_hits stays high
  • expired_keys spikes just before the database spike
  • Miss rate is concentrated on a few keys or one prefix
In the application
  • Spikes line up with deploys, restarts, or warm-up jobs
  • Connection pool waits rise on endpoints that read the hot key
  • Traces show many parallel spans running the same loader

The correlation step is the same one used for any incident: line up the metric spike with the code and deploy timeline. How to Correlate Logs, Traces, and a Code Change walks through it, and Which Commit Caused the Production Incident? covers narrowing it to a change.

TechniquePreventsDoes not preventMain caveat
TTL jitterMass synchronized expiryA single hot key stampedeMust be in the shared helper
Single-flightDuplicate loads within a processOne query per instanceLeader errors and timeouts are shared
Redis SET NX PX lockDuplicate loads across instancesWaiting, if there is no stale copyPX sizing; not a correctness lock
Soft TTL / SWRThe miss window on hot keysCold keys and cold startsReaders see bounded staleness
XFetchExpiry-time stampedes, lock-freeCold starts; rare double refreshNeeds Δ stored with the value
WarmingCold-start storms after deploysKeys outside the warmed setWarm-up itself needs jitter
Negative cachingRepeated misses on absent IDsMisses on real keysShort TTL; clear on create

The pull requests that cause cache stampedes

Stampedes rarely come from a new cache. They come from small changes to an existing one, reviewed as refactors. Each looks harmless in a diff and only becomes dangerous under concurrency and real traffic:

PR: “Warm product cache on startup”synchronized expiry
@@ def on_startup():
+   for product in db.top_products(limit=10_000):
+       r.set(f"product:{product.id}", json.dumps(product), ex=3600)
    # every pod runs this at deploy time: 10,000 keys, one TTL,
    # all expiring together one hour after the deploy
PR: “Simplify cache invalidation”removes the stale copy
@@ def update_price(product_id, price):
    db.update_price(product_id, price)
-   refresh_async(f"product:{product_id}")   # stale served during refresh
+   r.delete(f"product:{product_id}")        # next read misses
    # on a hot product with frequent price updates, every write
    # now opens a miss window, and nothing coalesces the reload

Delete-on-write is a legitimate, often recommended invalidation strategy: it avoids races where an older value overwrites a newer one. The risk is removing it from a path that previously served stale data without adding coalescing on the read side. The same review question applies to the others:

None of these is visible as a bug in the changed lines. The risk comes from how hot the key is, how expensive the loader is, and how many instances run the code, which is why blast radius matters more than diff size. How to Review a Pull Request for Production Reliability Risks covers the general review method; caching is one of the categories where it pays off most.

How Tomosu helps

Tomosu analyzes the standing codebase and each pull request for production reliability risk, and cache expiry logic is a case where the risk depends on context the diff does not show. For cache stampedes, Tomosu flags and maps:

These findings feed the Production Reliability Index alongside the Fragility, Drift, and Code Volatility signals, so a warm-up job with a fixed TTL is surfaced in review, not an hour after the deploy.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

What is a cache stampede?

A cache stampede happens when a frequently read cache key expires or is deleted and many concurrent requests miss it at the same time. Each request then recomputes the same value, usually by running the same database query, so the database receives a burst of identical work roughly equal to the request rate multiplied by the recompute time. It is also called a cache miss storm, dogpile effect, or thundering herd.

How do you prevent a cache stampede?

Make sure only one caller recomputes a value and that other callers do not wait on an empty cache. Use single-flight request coalescing in each process, a short Redis lock with SET NX PX across instances, TTL jitter so keys written together do not expire together, and a soft TTL with stale-while-revalidate or probabilistic early expiration so the value is refreshed before it disappears. Warm caches before traffic and avoid flushing shared caches during deploys.

What is the difference between a cache stampede and a cache miss storm?

They describe the same failure. Cache stampede usually refers to many requests missing one hot key at once. Cache miss storm is often used for the wider case where many keys miss together, for example after a deploy, a cache flush, or a warm-up job that gave thousands of keys the same TTL. Both overload the backend with work the cache normally absorbs.

How do I prevent a Redis cache stampede with a lock?

On a miss, try SET lock:key token NX PX 5000. The caller that gets OK recomputes the value, writes it to the cache, and releases the lock with a compare-and-delete script that only deletes the lock if the token still matches. Callers that do not get the lock should serve a stale copy, or wait with jittered sleeps and re-read the cache, rather than querying the database. Set PX from measured recompute time.

Does TTL jitter prevent a cache stampede?

Jitter prevents mass synchronized expiry, where many keys written at the same time expire at the same time. It does not protect a single hot key, which still expires at some instant and can still stampede. Combine jitter with request coalescing or a soft TTL for hot keys.

What is probabilistic early expiration or XFetch?

XFetch, from the 2015 VLDB paper Optimal Probabilistic Cache Stampede Prevention by Vattani, Chierichetti, and Lowenstein, lets each reader decide at random to recompute a value before it expires. The chance rises as expiry approaches and is scaled by how long the last recompute took, so one reader usually refreshes a hot key early while others keep reading the cached value. It needs no lock.

Is a Redis SET NX lock safe enough for cache stampede protection?

Yes for this purpose, because the lock only limits duplicate work. If the lock expires early or is lost in a Redis failover, the worst case is a second recompute. Do not rely on the same single-instance lock for operations that must happen exactly once, such as charging a payment, where you need fencing or a stronger coordination mechanism.

Why does my database spike after every deploy even though the cache hit ratio is high?

Deploys often empty caches: in-process caches start cold in every new instance, and deploy scripts or migrations may flush Redis or change the key prefix. The first requests after the deploy all miss together. Warm the hot set before the instance receives traffic, roll out gradually, and avoid flushes or prefix changes as part of a routine deploy.


This is the last post in the Production Debugging series. Pool timeouts, OOMKills, retry storms, duplicate webhooks, and cache stampedes share one shape: code that is correct for a single request and wrong under concurrency, shipped in a change that looked small. Tomosu maps those paths before they reach production. Assess your repository →