A retry storm in microservices rarely starts with a bad idea. It starts with a reasonable one: “this call is flaky, let’s retry it.” The pull request is three lines long, the tests pass, and the reviewer approves. What nobody sees in the diff is that the gateway above and the HTTP client below already retry the same call. The new line turns one failing request into 27.
A retry storm is a feedback loop in which retries add load to a service that is already failing because of load. When several layers each retry, the attempts multiply: three layers with three attempts each send up to 27 requests to the bottom service for one user action. That is how retries turn a brief slowdown into a cascading failure.
- Retry at one layer only, usually the one closest to the caller that knows the deadline.
- Use capped exponential backoff with full jitter, never a fixed interval.
- Cap retries with a budget (for example, about 10% of requests) and a small attempt limit.
- Do not retry non-idempotent operations without an idempotency key, client errors, or explicit overload signals.
- Propagate deadlines so no layer retries work its caller has already abandoned.
This guide explains the mechanism, shows the pull request pattern that most often creates retry amplification, and gives a review checklist you can apply before merge. The core idea comes from Google’s SRE book chapter Addressing Cascading Failures, which treats retries as one of the main ways overload spreads from one service to the next.
What is a retry storm in microservices?
A retry storm happens when a dependency slows down or starts failing, its callers respond by retrying, and the retries push the dependency further past its capacity. The dependency does not get a chance to recover, because the traffic it receives is now larger than the traffic that broke it.
Three properties make it different from an ordinary incident:
- It is triggered by something small. A garbage collection pause, a slow deploy, a hot partition, or losing one zone briefly raises the error rate. On its own, that would clear in seconds.
- It is sustained by the callers, not the original fault. Once the retries start, removing the trigger does not end the incident. The load is now self-generated.
- It is invisible in any single diff. Each retry was added by a different team, in a different file, at a different time, and each looks reasonable in isolation.
That last property is why retry storms belong in a pre-merge review. The information needed to see the problem exists in the repository and its configuration. It is spread across layers that a line-by-line diff review does not connect. We have written before about why code review and production reliability are different questions; retries are one of the clearest examples.
Why do nested retries multiply instead of add?
Each layer that retries treats the layer below it as a black box. If the edge gateway makes up to 3 attempts, and each of those reaches a service that makes up to 3 attempts, and each of those uses an HTTP client that makes up to 3 attempts, the bottom service can receive 3 × 3 × 3 = 27 requests for one user action. The SRE book gives the same arithmetic with four attempts per layer across three layers: 43 = 64 attempts on the database.
Two details make this worse than the arithmetic suggests:
- The worst case is the common case during an incident. When the bottom service is healthy, retries are rare and the multiplier is close to 1. When it is overloaded, nearly every attempt fails, so every layer exhausts its attempts. Amplification is highest precisely when capacity is lowest.
- Abandoned attempts keep running. If the gateway times out and retries while the inner layers are still working through their own retries, the new chain overlaps the old one. The bottom service now serves two chains for one user, and the first chain’s answer goes nowhere.
The pull request that adds the third retry layer
Here is the kind of change that creates the problem. Checkout calls inventory to reserve stock. Inventory has had a few 503s during deploys, so a developer adds a retry. The diff looks like good engineering:
import requests
+from tenacity import retry, stop_after_attempt, wait_fixed, retry_if_exception_type
from platform.http import session
+@retry(stop=stop_after_attempt(3), wait=wait_fixed(0.5),
+ retry=retry_if_exception_type(requests.RequestException))
def reserve_stock(order_id, items):
resp = session.post(
f"{INVENTORY_URL}/reservations",
json={"order_id": order_id, "items": items},
timeout=1.0,
)
resp.raise_for_status()
return resp.json()
Nothing in these lines is wrong in isolation. The problem is what the diff does not show. The shared session, imported from a platform package, already has retries mounted on it:
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
_retry = Retry(
total=2, # 2 retries = 3 attempts
backoff_factor=0.2,
status_forcelist=[502, 503, 504],
allowed_methods=None, # None = retry every verb, including POST
)
session = requests.Session()
session.mount("https://", HTTPAdapter(max_retries=_retry))
And the route in front of checkout, owned by the platform team, retries server errors too:
route:
cluster: checkout
timeout: 6s
retry_policy:
retry_on: "5xx,reset,connect-failure"
num_retries: 2 # 3 attempts
per_try_timeout: 2s
Read together, the three files describe a much riskier change than the diff suggests:
| Issue | Where it comes from | Why it matters | What a reviewer should ask |
|---|---|---|---|
| 27× amplification | Gateway × decorator × client | A short inventory brownout becomes a sustained overload | Which single layer should own retries for this path? |
| Retrying a non-idempotent POST | allowed_methods=None plus the decorator | A timed-out reservation may have succeeded; retrying can reserve stock twice | Does /reservations accept an idempotency key? |
| Retrying client errors | raise_for_status() raises HTTPError, a RequestException | A 409 “out of stock” or 400 is retried three times and can never succeed | Which status codes are actually transient? |
| Synchronized retries | wait_fixed(0.5) | Every failed caller retries at the same moment, in waves | Is the backoff exponential, capped, and jittered? |
| Retries past the deadline | Inner retries can take ~10 s; gateway per-try timeout is 2 s | Work continues for callers that already gave up, and overlaps the gateway’s own retry | Is the remaining deadline checked before each attempt? |
This is a textbook case of a change whose risk lives outside its lines. Assessing it means following the call path and its configuration, the same way you would assess the blast radius of a code change.
How do retries cause cascading failures?
A cascading failure is a failure that spreads because the system’s response to it creates more of the same failure. Retries are a direct path to that. They convert errors into load, and in an overload event, load is the cause of the errors.
Several things make the loop hard to escape once it starts:
- Recovery needs headroom. A service that comes back from a restart has cold caches and fresh connection pools. It is slower than normal, just when it receives the most traffic.
- Timeouts shrink effective capacity. When a request times out at the caller, the work the server did for it is wasted. The server spends capacity on answers nobody reads.
- Overload spreads upward. Callers that wait on retries hold threads, connections, and memory. Their own capacity drops, and their callers start retrying them. If those waits happen inside a database transaction, you also get the pool starvation described in HikariCP: Connection Is Not Available, Request Timed Out.
The same shape shows up with caches: many callers acting at once on the same signal, turning a small event into a load spike on a shared dependency. See Cache Stampede: Why a Healthy Cache Can Overload Your Database for that variant.
A retry is a bet that the failure was random. During overload, the failure is not random, and every retry makes the bet worse.
What should never be retried?
Retries are safe only when two things are true: the failure is likely to be transient, and repeating the operation cannot cause a different outcome. Most retry storms and many data bugs come from ignoring one of those conditions.
| Response | Transient? | Retry? | Notes |
|---|---|---|---|
| Connection refused, DNS failure before sending | Often | Yes, with backoff | The request never reached the server, so repeating it is safe even for writes |
400, 401, 403, 404, 409, 422 | No | No | The same request will fail the same way. Retrying only adds load |
429 Too Many Requests | Yes, but it is an overload signal | Only after Retry-After, within budget | The server is explicitly asking for less traffic |
502, 503, 504 | Usually | At one layer, with jittered backoff and budget | Treat a 503 that means “overloaded” as a signal to back off, not try again |
500 | Unknown | Only for idempotent calls | Often a deterministic bug; retries rarely help |
| Timeout after the request was sent | Unknown | Only if idempotent or keyed | The server may have completed the work. The outcome is unknown, not failed |
Non-idempotent operations
A POST /reservations that timed out may have reserved the stock. A retry reserves it again. The same applies to charging a card, sending an email, or publishing a message. Before any write is retried, it needs an idempotency key that the server stores and checks, so a repeated request returns the original result instead of repeating the effect. The mechanics are the same as deduplicating webhooks; see How to Prevent Duplicate Webhook Processing.
Many HTTP client libraries retry only idempotent methods by default for this reason. In urllib3, the default allowed_methods set excludes POST; setting it to None, as the shared client above does, turns that protection off for every caller of that session.
Overload errors
Some errors mean “I am broken right now,” and others mean “I am too busy.” Retrying the first kind may help. Retrying the second kind is the retry storm. The SRE book’s chapter on Handling Overload describes a dedicated “overloaded; don’t retry” error that a backend returns once its own retries are exhausted, so the layers above do not multiply them. If your services cannot tell those cases apart, add the distinction: a specific status, header, or error code that callers must not retry.
Exponential backoff with jitter: why fixed delays synchronize retries
A fixed delay, like wait_fixed(0.5) in the PR above, means every caller that failed at the same moment retries at the same moment. The failures came in together, so the retries arrive together too, as waves that hit the dependency every half second. Exponential backoff spreads retries out over longer intervals. Jitter spreads them out across callers.
The AWS Architecture Blog post Exponential Backoff And Jitter compares several strategies with a simulation. The one most teams should default to is full jitter:
sleep = random_between(0, min(cap, base * 2 ** attempt))
# base = 100 ms, cap = 1 s:
# attempt 1: random in [0, 200 ms]
# attempt 2: random in [0, 400 ms]
# attempt 3: random in [0, 800 ms]
# attempt 4+: random in [0, 1 s]
In Python, tenacity provides this as wait_random_exponential(multiplier=..., max=...). gRPC’s built-in retry policy applies randomized exponential backoff through initialBackoff, maxBackoff, and backoffMultiplier. Whatever the library, check three things in review: the delay grows, it is capped, and it is randomized.
Backoff and jitter shape when retries arrive. They do not limit how many there are. With 27 attempts per request, perfectly jittered retries are still 27 times the load. You need the limits in the next section too.
Retry budgets, deadlines, circuit breakers, and load shedding
Five controls, used together, keep retries from turning into a cascading failure. Each one answers a different question.
1. Retry at one layer only
Pick one layer per call path to own retries, usually the layer closest to the caller that knows the user’s deadline and whether the operation is safe to repeat. Every other layer passes failures through. This single rule changes the worst case from a product to a sum. If the gateway must retry, restrict it to failures where the request never reached the upstream (connection failures, refused streams), and disable retries in the libraries below.
2. Retry budgets
An attempt limit (“at most 3 tries”) bounds retries per request. A retry budget bounds retries per client or process: retries are allowed only while they stay under a fixed fraction of normal requests. The SRE book describes a per-request limit of three attempts combined with a per-client budget of about 10% of requests. When most requests are failing, the budget runs out quickly and the client stops retrying, which is the behavior you want during an overload. gRPC offers this as retryThrottling (maxTokens, tokenRatio) and Envoy offers a retry_budget on the cluster’s circuit breaker thresholds.
3. Deadlines and timeout propagation
Every request should carry a deadline, and every layer should check the remaining time before starting an attempt. A retry that cannot finish before the caller gives up is pure waste. gRPC propagates deadlines natively; over HTTP, pass the remaining budget in a header your services agree on, and derive each call’s timeout from it.
4. Circuit breakers
A circuit breaker watches the error rate for a dependency and, above a threshold, stops sending requests for a cool-down period, then lets a few probes through. A retry budget limits the extra load retries add; a circuit breaker removes most of the load entirely while the dependency is unhealthy, and fails fast so callers do not hold threads waiting. Resilience4j and Polly provide circuit breakers as library components; at the proxy layer, Envoy offers circuit breaking thresholds and outlier detection. Review that the breaker’s open state produces a clear, non-retryable error, or the layer above will retry the breaker.
5. Load shedding on the server
The callee protects itself as well. When a service is near capacity, it should reject excess requests cheaply and early, before doing expensive work, and return a response that tells callers not to retry immediately. Rejecting 20% of requests quickly keeps the other 80% fast. Accepting all of them makes all of them slow, which triggers timeouts and retries everywhere upstream.
The same change, done safely
Here is the checkout call rewritten so that retries live in exactly one place, with jittered backoff, a budget, a deadline, and an idempotency key. First, a small shared helper:
import random, threading, time
RETRYABLE_STATUS = {502, 503, 504} # never 4xx; 429 needs Retry-After handling
class RetryBudget:
"""Token bucket: each request earns `ratio` tokens, each retry spends 1."""
def __init__(self, ratio=0.1, max_tokens=10.0):
self.ratio, self.max_tokens = ratio, max_tokens
self.tokens = max_tokens
self._lock = threading.Lock()
def on_request(self):
with self._lock:
self.tokens = min(self.max_tokens, self.tokens + self.ratio)
def try_spend(self):
with self._lock:
if self.tokens >= 1.0:
self.tokens -= 1.0
return True
return False # budget empty: stop retrying
def call_with_retries(send, *, deadline, budget, max_attempts=3, base=0.1, cap=1.0):
budget.on_request()
attempt = 0
while True:
remaining = deadline - time.monotonic()
if remaining <= 0:
raise TimeoutError("deadline exceeded before attempt")
resp = send(timeout=min(1.0, remaining)) # never outlive the caller
attempt += 1
if resp.status_code not in RETRYABLE_STATUS:
return resp # success or a non-transient error
if attempt >= max_attempts or not budget.try_spend():
return resp # let the caller see the failure
delay = random.uniform(0, min(cap, base * 2 ** attempt)) # full jitter
if time.monotonic() + delay >= deadline:
return resp
time.sleep(delay)
from platform.http import no_retry_session # Retry(total=0): no hidden layer
from platform.retries import RetryBudget, call_with_retries
_inventory_budget = RetryBudget(ratio=0.1)
def reserve_stock(order_id, items, *, deadline):
def send(timeout):
return no_retry_session.post(
f"{INVENTORY_URL}/reservations",
json={"order_id": order_id, "items": items},
headers={"Idempotency-Key": f"reserve-{order_id}"}, # safe to repeat
timeout=timeout,
)
resp = call_with_retries(send, deadline=deadline, budget=_inventory_budget)
resp.raise_for_status()
return resp.json()
The sketch handles HTTP status codes only. A production version would also retry connection errors that happen before the request is sent, honor Retry-After on 429 and 503, and emit a metric for every retry and every budget refusal. Retries that are not measured are how a storm goes unnoticed until it is an outage.
And the gateway stops retrying server errors on this route, keeping only failures where the request never reached checkout, bounded by a budget:
route:
cluster: checkout
timeout: 3s
retry_policy:
retry_on: "connect-failure,refused-stream" # request never reached checkout
num_retries: 1
# in the checkout cluster definition:
circuit_breakers:
thresholds:
- priority: DEFAULT
retry_budget:
budget_percent: { value: 10.0 }
min_retry_concurrency: 3
The worst case for one user request is now roughly 2 × 3 = 6 attempts on inventory, and far fewer in practice, because the budgets drain as soon as failures become widespread. More importantly, retrying a reservation can no longer reserve stock twice.
Retry storm checklist for pull request review
Use this whenever a pull request adds or changes a retry, a timeout, a client library, or gateway configuration. The first group is the one reviewers most often skip, because answering it requires looking outside the diff.
- List every layer that retries this call: client SDK, gateway or mesh, service code, HTTP or gRPC client, driver
- Multiply the attempt counts; flag anything above single digits
- Name the one layer that should own retries
- Only transient errors; no 4xx
- No retry on explicit overload signals
- Writes carry an idempotency key the server enforces
- Library defaults for methods are not widened
- Exponential, capped, and jittered backoff
- No fixed-interval or immediate retries
Retry-Afteris honored when present
- Small attempt limit plus a retry budget
- Remaining deadline checked before each attempt
- Per-attempt timeout shorter than the caller’s timeout
- Circuit breaker errors are not retried upstream
- Retries and budget refusals are emitted as metrics
A PR that fails the first group should be treated as a high-risk change regardless of its size. It fits the pattern in How to Identify High Risk Pull Requests: a small diff that changes behavior on a shared, critical path. For the wider review practice, see How to Review a Pull Request for Production Reliability Risks.
How Tomosu helps spot a retry storm before merge
A style-focused review of PR #4812 has little to object to. The decorator is idiomatic, the imports are tidy, the function is short. The risk only appears when the diff is read together with the shared client it calls through and the gateway route in front of it. That cross-layer reading is what Tomosu does on a pull request:
- Maps the call path the change touches, from the entry route through service code to the outbound client, and identifies every layer that retries along it, including client and gateway configuration in the repository that the diff does not touch.
- Surfaces the combined attempt count as evidence for the reviewer: which layers retry, how many attempts each allows, and what the worst case is when the dependency is failing.
- Flags the specific risk patterns described above: retries on non-idempotent writes without a key, retries on 4xx or overload responses, fixed or unjittered delays, and retries with no deadline check.
- Weighs the finding by blast radius, so a new retry layer on the checkout path is ranked above the same pattern in an internal batch job.
PR #4812 adds a retry layer to POST {INVENTORY_URL}/reservations
layer 1 deploy/gateway/checkout-route.yaml num_retries: 2 3 attempts
layer 2 checkout/inventory_client.py stop_after_attempt 3 attempts (new)
layer 3 platform/http.py Retry(total=2) 3 attempts
worst case per user request: 27 calls to inventory-service
also: POST retried without Idempotency-Key; 4xx retried; wait_fixed (no jitter)
evidence to check: inventory 503 rate during deploys, retry metrics, deadline
These findings feed the Fragility Index and the overall Production Reliability Index, so a three-line retry decorator is weighed by what it does to the call path, not by its line count. It fits into the broader practice of pre merge reliability analysis.
Scan your repository with Tomosu →
Key takeaways
- A retry storm is a feedback loop: retries add load to a service that is failing because of load.
- Retries multiply across layers. Three layers of three attempts is 27 requests at the bottom, and the worst case is the incident case.
- Retry at one layer only. Everywhere else, pass the failure through.
- Never retry client errors, explicit overload signals, or non-idempotent writes without an idempotency key.
- Use capped exponential backoff with full jitter, and bound retries with a small attempt limit plus a retry budget.
- Propagate deadlines, and add circuit breakers and load shedding so an unhealthy dependency gets less traffic, not more.
- The risk lives outside the diff. Review the call path and its configuration, not just the changed lines.
Frequently asked questions
What is a retry storm?
A retry storm is a feedback loop in which clients retry failed requests to a service that is failing because it is overloaded. The retries add more load, which causes more failures, which cause more retries. It often starts from a brief slowdown and continues after the original trigger is gone, because the callers now generate the excess traffic.
How do retries cause cascading failures in microservices?
When several layers each retry, attempts multiply: three layers with three attempts each can send 27 requests to the bottom service for one user action. During an overload nearly every attempt fails, so every layer uses its full allowance. Callers waiting on retries also hold threads and connections, lose capacity themselves, and get retried by their own callers, so the overload spreads upward.
How many retries is too many?
Count attempts across the whole call path, not per layer. A common guideline is a small number of attempts, such as three, at a single layer, combined with a retry budget that keeps retries to roughly 10% of normal requests. If multiplying the attempt counts of every retrying layer on a path gives more than single digits, the path needs to be simplified.
What is exponential backoff with jitter?
Exponential backoff increases the delay between retries, typically doubling it each attempt up to a cap. Jitter randomizes each delay so that callers that failed together do not retry together. With full jitter, the delay is a random value between zero and min(cap, base × 2^attempt). Backoff and jitter control when retries happen, not how many, so they should be combined with attempt limits and a retry budget.
Should I retry on HTTP 429 or 503?
Treat both as possible overload signals. For 429, wait at least as long as the Retry-After header says, and only retry within your retry budget. For 503, retry at one layer with jittered backoff if the error is transient, but do not retry if the service indicates it is overloaded. Never retry 4xx errors such as 400, 404, or 409, because the same request will fail the same way.
Is it safe to retry a POST request?
Only if the operation is idempotent or carries an idempotency key that the server stores and checks. A POST that timed out may have succeeded on the server, so a blind retry can create a duplicate order, reservation, or charge. It is safe to retry a POST when the connection failed before the request was sent, because the server never received it.
What is the difference between a retry budget and a circuit breaker?
A retry budget limits how many retries a client may send, usually as a fraction of its normal requests, so retries cannot multiply load during an outage. A circuit breaker stops sending requests to an unhealthy dependency for a period and fails fast instead. The budget caps the extra load from retries; the breaker removes most of the load entirely while the dependency recovers. They work best together.
Where should retries live in a microservice call chain?
At one layer per call path, usually the one closest to the caller that knows the request deadline and whether the operation is safe to repeat. Other layers should pass failures through instead of retrying. If a gateway or service mesh retries at all, limit it to failures where the request never reached the upstream, and disable retries in client libraries below the owning layer.
A retry storm is written one reasonable line at a time, in different files, by different teams. Tomosu reads those layers together, so the multiplication is visible before merge instead of during the incident. Assess your repository →