The release went out, the dashboards are green, and a customer reports that the price they changed ten minutes ago is still the old one. Nothing is erroring. Somewhere between the database and the user, a copy of the data stopped being updated, and the release is the reason.
Stale data after a release almost always means a cache is answering with a copy that the new code no longer invalidates. First find which layer served the old value (CDN, proxy, in-process cache, Redis, ORM cache, or a read replica). Then find which change in the release broke the invalidation for that layer.
- New write path: a job, bulk endpoint, or migration updates rows but never evicts.
- Key drift: readers and writers now build different cache keys.
- Ordering race: the evict runs before the database commit.
- Forgotten layer: a new
Cache-Controlheader, per-pod cache, or replica read.
Flushing the cache makes the symptom disappear for a while, which is why stale data bugs tend to come back. This guide shows how to find the layer that answered, how to tie it to a specific change in the release, and how to fix the invalidation path so it stays fixed.
What does stale data after a release actually mean?
Stale data is a response that reflects an older state than the latest committed write, by more than the staleness you agreed to accept. Every cache is stale for some window by design. A 60-second TTL on a product listing is a decision. A profile page that shows the old email address for six hours after the user changed it is a bug.
Before debugging the cache, rule out two look-alikes:
- Wrong data, not old data. If the value never existed in the database, the problem is computation or serialization, not invalidation. Compare the stale value with the row’s history or audit log.
- The write never happened. If the primary database still has the old value, the write failed or went somewhere else. No cache is involved.
If the primary has the new value and the user sees the old one, a copy is involved. The question is which one.
Which cache layer served the stale response?
A single request can pass through six or more copies of the same data before it reaches the source of truth. Each layer has its own invalidation mechanism, its own way of going stale, and its own evidence. Most teams think of Redis first, but after a release the stale copy is just as often in a layer nobody thought of as a cache.
The HTTP layers are the easiest to check, because they announce themselves in response headers. The Age header (RFC 9111) is the number of seconds the response has spent in a cache. Many CDNs also add a status header, such as X-Cache on CloudFront and Fastly or cf-cache-status on Cloudflare.
# Age > 0, or HIT in a cache status header, means a cache answered
curl -sS -D - -o /dev/null https://api.example.com/v1/products/42 \
| grep -iE '^(age|cache-control|x-cache|cf-cache-status|etag|last-modified):'
# Ask the origin directly, skipping the CDN (internal host or port-forward)
curl -sS http://products.internal:8080/v1/products/42 | jq '.price, .updated_at'
| What you observe | What it means |
|---|---|
Edge response has Age > 0 and is stale; origin is fresh | An HTTP cache (CDN or proxy) is serving an unpurged copy. Check what the release did to Cache-Control and purge rules. |
| Origin is stale; primary database is fresh | An application cache or replica is involved. Go to the in-process cache, Redis, ORM cache, and read routing. |
| Refreshing flips between old and new values | Different instances hold different copies: per-pod caches, several edge locations, or a mixed-version rollout. |
| Stale for a few seconds after a write, then fresh | Usually replica lag or a short TTL working as designed. Check whether the read should have gone to the primary. |
| Stale until someone restarts the service | An in-process cache with no TTL and no cross-instance eviction. |
| Stale until the TTL expires, then fresh | The evict never ran or deleted the wrong key. The TTL is doing all the work. |
| Stale indefinitely, even after TTL should have passed | The entry has no TTL (TTL returns -1 in Redis), or a background refresher keeps rewriting the old value. |
For the shared cache, ask Redis directly what it holds. Use --scan (which uses SCAN) rather than KEYS on a production instance:
# Which key versions exist for this record?
redis-cli --scan --pattern 'product:*42'
# -1 = no expiry set, -2 = key does not exist
redis-cli TTL product:v2:42
# Compare the cached value with the row in the primary
redis-cli GET product:v2:42 | jq '.updated_at'
Two keys for one record (product:42 and product:v2:42) are a strong hint that the release changed the key format. A TTL of -1 means a missed eviction will never heal on its own.
Why do releases break cache invalidation?
Cache invalidation is a contract spread across many files: every code path that writes a record must evict or update every copy of it, using the same key the readers use, after the write is durable. A release rarely breaks the cache itself. It adds or changes one side of that contract without the other.
1. A new write path that skips the evict
The original update endpoint evicts correctly. The release adds a CSV import, an admin bulk edit, a backfill migration, or a queue consumer that updates the same rows through a different path. That path writes to the database and never touches the cache. This is the most common cause, and code review misses it because the diff contains only the new path. The evict that is missing is in a file nobody changed.
In Spring, the same gap appears in a quieter form. Caching annotations work through a proxy, so a call from one method to another in the same class does not go through the proxy and @CacheEvict never runs:
@Service
public class ProductService {
@CacheEvict(cacheNames = "products", key = "#id")
public void updatePrice(long id, BigDecimal price) {
repo.updatePrice(id, price);
}
// Added in this release: bulk repricing from a CSV import
@Transactional
public void applyPriceList(List<PriceRow> rows) {
for (PriceRow r : rows) {
this.updatePrice(r.id(), r.price()); // self-call: no proxy, no evict
}
}
}
The Spring cache abstraction documentation calls this out: in the default proxy mode, only external method calls coming in through the proxy are intercepted. The unit test for applyPriceList passes because it checks the database, not the cache.
2. Key drift: readers and writers build different keys
The release changes the cached shape (a new field, a different serializer, a currency) and bumps the key to product:v2:{id} so old entries are not misread. The readers were updated. The eviction in the admin service was not, so it still deletes product:{id}, a key nobody reads anymore. Every write now leaves the v2 entry untouched until its TTL expires.
// catalog/read.ts (changed in this release)
const key = `product:v2:${id}`; // new shape adds currency
const hit = await redis.get(key);
// admin/prices.ts (unchanged)
await db.query('UPDATE products SET price = $1 WHERE id = $2', [price, id]);
await redis.del(`product:${id}`); // deletes the v1 key; readers use v2
// cache/keys.ts: the only place product keys are built
export const productKey = (id: string) => `product:v2:${id}`;
export const allProductKeys = (id: string) => [
`product:v2:${id}`,
`product:${id}`, // v1: remove once no running version reads it
];
// admin/prices.ts
await db.query('UPDATE products SET price = $1 WHERE id = $2', [price, id]);
await redis.del(allProductKeys(id)); // evicts every live version
3. Eviction moved before the commit
A refactor wraps an existing method in a transaction, or moves the evict into a helper that runs earlier. The eviction still happens, but now it can happen before the new value is visible to other connections. That opens a race covered in the next section.
4. Serialization changes that read old entries wrong
The release adds a field to a cached object without changing the key. Old entries still deserialize, but the new field comes back empty or with a default. The page shows “no discount” or a missing address, which looks exactly like stale data. If entries are not versioned, every shape change is an invalidation event for the whole keyspace.
5. HTTP caching headers that changed
A framework upgrade or middleware change starts sending Cache-Control: public, max-age=300 on an endpoint that used to send nothing or no-store. Now the CDN and browsers keep a five-minute copy that no purge covers. If the response is per-user and the cache key ignores the auth header, this is worse than stale data: one user’s response can be served to another. Per-user responses should carry private or no-store.
6. Per-pod caches and read replicas
An in-process cache (a Map, an LRU, Caffeine, Guava) lives inside one instance. When one pod handles the write and evicts its own copy, the other pods keep theirs until their TTL, or forever. Similarly, a release that routes a read to a replica turns replication lag into user-visible stale data, especially for read-your-own-writes flows. Neither shows up in a single-instance test environment.
What is the evict-before-commit race?
The evict-before-commit race happens when a writer deletes the cache entry before its database transaction commits, and a concurrent reader re-caches the old row in between. The writer did everything “right”: it evicted. But the reader saw an empty cache, went to the database, got the pre-commit value, and wrote it back. The old value then sits in the cache until the TTL expires.
This race is easy to introduce without touching the cache code. In Spring, @CacheEvict runs after its method returns by default (beforeInvocation = false). If that method is called from inside a larger @Transactional method, “after the method” is still before the outer commit. Adding a transaction around existing code is enough to move every eviction inside it into the race window.
@Transactional
public void applyPriceList(List<PriceRow> rows) {
for (PriceRow r : rows) repo.updatePrice(r.id(), r.price());
List<Long> ids = rows.stream().map(PriceRow::id).toList();
events.publishEvent(new PricesChanged(ids));
}
@TransactionalEventListener // default phase: AFTER_COMMIT
public void onPricesChanged(PricesChanged e) {
Cache cache = cacheManager.getCache("products");
e.ids().forEach(cache::evict); // runs only once the new rows are visible
}
This fixes the self-invocation gap and the ordering together. It still leaves one narrow window: a reader that fetched the old row just before the commit can write it to the cache just after the evict. Closing that completely takes a version check on write (only set the cache if the value is newer than what is there) or a short TTL that bounds the damage. For most data, a TTL sized to your staleness tolerance is the pragmatic answer.
Writing the new value into the cache from the writer (update on write) looks fresher than deleting it, but two concurrent writers can finish in the opposite order to their commits and leave the older value cached. Delete on write, after commit, and let the next reader repopulate, is the safer default.
How do rolling deploys and rollbacks cause stale data?
During a rolling deploy, old and new versions of the service run side by side for minutes. If the release changed a cache key, each version evicts only the key it knows. Writes handled by new pods leave the old key untouched, and old pods keep serving it. The problem also runs the other way, which is why a rollback does not always clear it.
The safe order mirrors an expand and contract database migration. First ship a release whose writers evict both the old and new keys, while readers still use the old one. Then ship the release whose readers switch to the new key. Only when no running or rollback-candidate version reads the old key, remove it from the eviction list. If you are weighing whether to roll back at all, this leftover-entry problem is one of the hazards covered in How to Decide Whether to Roll Back a Release.
A cache is a second copy of your data with its own write path. Every release that adds a write without touching that path ships stale data.
How do you debug stale data after a release, step by step?
Work from a concrete stale sample, not a general report. The goal is to name the layer that answered and the change that stopped invalidating it.
- Capture one stale response. Record the full response headers, the request ID, the time, the instance or pod that served it, and the value you expected. Without the instance, per-pod caches and mixed versions are invisible.
- Ask the source of truth. Read the record from the primary database, then from the origin service with every cache bypassed. If the primary is also old, stop: this is a write problem, not a cache problem.
- Walk the layers from the outside in. Check
Ageand cache status headers for the browser, CDN, and proxy. Then check the in-process cache, the shared cache, the ORM cache, and whether the read went to a replica. - Compare the key read with the key evicted. Log the exact key the reader used and the key the writer deleted. List which key versions exist in Redis for that record, and their TTLs.
- Diff the release for invalidation changes. Look for new write paths to the same tables, changed key builders, changed serialization, changed
Cache-Controlheaders, new read routing, and evictions that moved inside a transaction. - Confirm with a targeted reproduction. Write the record and immediately read it through two different instances, during a rollout if key versions are involved. Then fix the invalidation path, and evict only the affected keys.
The timing matters as much as the code. If staleness started exactly when the deploy began, the rollout itself (mixed versions) is suspect. If it started when a new job first ran, the job is. Correlating logs, traces, and the code change on one timeline usually separates the two in minutes, and finding the commit that caused the incident narrows a large release to the change that touched a write path.
How do you fix stale cache data without turning the cache off?
| Cause | Fix | What not to do |
|---|---|---|
| New write path skips the evict | Route every write to the entity through one function that evicts after commit, or publish a change event that one listener handles. | Add an evict call to just the new path and leave the pattern in place. |
| Key drift | Build keys in one module. During a key change, evict every live version until no running version reads the old one. | Rename the key in readers only and rely on the TTL. |
| Evict before commit | Evict in an after-commit hook or from an outbox. Bound the remaining race with a TTL or a version check. | Evict twice with a sleep in between and call it fixed. |
| Serialization change | Version the key or embed a schema version in the value and treat mismatches as a miss. | Deploy a new shape into an unversioned keyspace. |
New Cache-Control header | Set caching headers explicitly per route. Use private or no-store for per-user data. Purge affected URLs. | Let a framework default decide what the CDN caches. |
| Per-pod in-process cache | Short TTL, or broadcast evictions (for example with Redis pub/sub), or move the data to the shared cache. | Restart pods as the invalidation mechanism. |
| Read moved to a replica | Send read-your-own-writes flows to the primary, or wait for the replica to catch up to the write position. | Treat replica lag as a cache bug and add evictions. |
Flushing the entire cache is tempting during an incident. On a busy service, it sends every read to the database at once, which can turn a correctness problem into an outage. Evict the affected keys or key prefix instead. Cache Stampede: Why a Healthy Cache Can Overload Your Database covers why a cold cache is dangerous and how to refill it safely.
What should you check in a PR that touches cached data?
Stale data bugs are cheapest to catch before merge, but only if the reviewer looks beyond the diff. The missing evict is usually in a file the PR did not change. Ask these questions of any change that writes to, reads from, or reshapes cached data:
- Does this PR add a write to a table or entity that is cached anywhere?
- Does that write go through the function that evicts?
- Does the evict run after the commit?
- Did the key format or the cached object change?
- Do all writers evict the new key and the old one?
- Is it safe during a rolling deploy and a rollback?
- Did any
Cache-Controlor CDN rule change? - Is there a new in-process cache or memoization?
- Did any read move to a replica?
Tests rarely catch these, because a single-instance test with an empty cache cannot show a stale copy. That is one of the escape classes in A PR Passed CI but Broke Production: What Did the Tests Miss?. For the wider review practice, see How to Review a Pull Request for Production Reliability Risks.
How Tomosu helps
Stale data after a release is a cross-file problem: the write is in one place, the cache read in another, and the eviction, if it exists, in a third. Tomosu analyzes the repository as a whole and each pull request against it, so it can surface the parts a diff hides:
- Cached reads and the writes that feed them: which code paths read an entity through a cache, and which paths write that entity.
- Write paths without a matching invalidation: a new job, endpoint, or consumer that updates cached data but never evicts, or evicts inside an open transaction.
- Key and shape changes: a PR that changes how a cache key is built or what is stored under it, with the other files that still use the old form.
- Blast radius: which endpoints serve the affected data, so a stale price on checkout is weighed differently from a stale avatar.
These findings feed the Production Reliability Index alongside other change risks. They are review signals with the evidence attached, not a replacement for checking the cache during an incident.
Scan your repository with Tomosu →
Key takeaways
- Stale data after a release means some copy of the data is no longer invalidated by the new code. First find which copy answered.
- Walk the layers outside in: browser, CDN, proxy, in-process cache, Redis, ORM cache, replica, primary.
Ageand cache status headers expose the HTTP layers. - The most common cause is a new write path that never evicts. The missing evict lives in a file the PR did not touch.
- Evict after the commit, not before. Keep a TTL as the upper bound on how long a missed eviction can hurt.
- A cache key change is a migration: evict every live version during the rollout, and remember that leftovers survive a rollback.
- Flush specific keys, not the whole cache, and fix the invalidation path so the stale data does not come back.
Frequently asked questions
Why do I see stale data after a deploy?
Usually because the release changed how data is written or cached without changing how the cache is invalidated. Common causes are a new write path that skips the evict, a cache key format change so readers and writers use different keys, eviction that now runs before the database commit, a new Cache-Control header, or a read moved to a lagging replica.
How do I know which cache served a stale response?
Work from the outside in. An Age header above zero or a HIT in X-Cache or cf-cache-status means an HTTP cache answered. If the origin is also stale, log the application cache key and hit or miss, check the key’s TTL and value in Redis, and compare with the primary database.
Should I delete the cache before or after the database write?
After the transaction commits. Deleting first lets a concurrent reader miss, read the old row, and put the old value back before the commit. Evicting after commit closes most of that window. A read that started before the commit can still write late, so keep a TTL or a version check as the upper bound on staleness.
Why does stale data appear on some requests but not others?
Something differs between requests: which pod served them (per-instance in-process caches), which CDN edge location answered, whether the read went to a lagging replica, or whether an old or new version of the service handled it during a rolling deploy. Record the instance, edge, and version with each stale sample.
Is flushing the whole cache a safe fix for stale data?
It is a blunt mitigation, not a fix. A full flush sends every read to the database at once and can cause a cache stampede on a busy service, and the stale entries come back as soon as the broken write path runs again. Evict the affected keys, then fix the invalidation path.
Can rolling back a release cause stale cache data?
Yes. If the release introduced a new key version, entries written under it survive the rollback, and the old code does not evict them. When the release is deployed again, it reads those leftover entries. Invalidate every live key version during a transition and give versioned keys a TTL.
How long should a cache TTL be to limit stale data?
Set the TTL to the longest staleness the business can accept if invalidation fails, not to a round number. The TTL is the safety net that bounds how long a missed eviction can hurt. Explicit invalidation on every write path is what keeps data fresh day to day.
A cache is a second copy of your data with its own write path. Tomosu flags write paths that skip it before merge, not after a customer sees an old price. Assess your repository →