Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Retry storms

Retries Can Cause an Outage: How to Spot a Retry Storm Before Merge

Tomosu AI·15 min read·

A retry storm in microservices rarely starts with a bad idea. It starts with a reasonable one: “this call is flaky, let’s retry it.” The pull request is three lines long, the tests pass, and the reviewer approves. What nobody sees in the diff is that the gateway above and the HTTP client below already retry the same call. The new line turns one failing request into 27.

Quick answer

A retry storm is a feedback loop in which retries add load to a service that is already failing because of load. When several layers each retry, the attempts multiply: three layers with three attempts each send up to 27 requests to the bottom service for one user action. That is how retries turn a brief slowdown into a cascading failure.

This guide explains the mechanism, shows the pull request pattern that most often creates retry amplification, and gives a review checklist you can apply before merge. The core idea comes from Google’s SRE book chapter Addressing Cascading Failures, which treats retries as one of the main ways overload spreads from one service to the next.

What is a retry storm in microservices?

A retry storm happens when a dependency slows down or starts failing, its callers respond by retrying, and the retries push the dependency further past its capacity. The dependency does not get a chance to recover, because the traffic it receives is now larger than the traffic that broke it.

Three properties make it different from an ordinary incident:

That last property is why retry storms belong in a pre-merge review. The information needed to see the problem exists in the repository and its configuration. It is spread across layers that a line-by-line diff review does not connect. We have written before about why code review and production reliability are different questions; retries are one of the clearest examples.

Why do nested retries multiply instead of add?

Each layer that retries treats the layer below it as a black box. If the edge gateway makes up to 3 attempts, and each of those reaches a service that makes up to 3 attempts, and each of those uses an HTTP client that makes up to 3 attempts, the bottom service can receive 3 × 3 × 3 = 27 requests for one user action. The SRE book gives the same arithmetic with four attempts per layer across three layers: 43 = 64 attempts on the database.

ONE USER REQUEST, THREE RETRYING LAYERS User request 1 call Edge gateway retry 3 attempts → 3 calls checkout @retry (this PR) 3 per call → 9 calls Shared HTTP client retry 3 per call → 27 calls 27 requests reach inventory-service for one user click Worst case = 3 × 3 × 3 = 27. One more retrying layer below, such as a DB driver, makes it 81.
Attempts multiply across layers. The worst case is exactly the case that matters: when the bottom service is failing, every attempt fails and every layer uses its full allowance.

Two details make this worse than the arithmetic suggests:

The pull request that adds the third retry layer

Here is the kind of change that creates the problem. Checkout calls inventory to reserve stock. Inventory has had a few 503s during deploys, so a developer adds a retry. The diff looks like good engineering:

checkout/inventory_client.py · PR #4812third retry layer
 import requests
+from tenacity import retry, stop_after_attempt, wait_fixed, retry_if_exception_type
 from platform.http import session

+@retry(stop=stop_after_attempt(3), wait=wait_fixed(0.5),
+       retry=retry_if_exception_type(requests.RequestException))
 def reserve_stock(order_id, items):
     resp = session.post(
         f"{INVENTORY_URL}/reservations",
         json={"order_id": order_id, "items": items},
         timeout=1.0,
     )
     resp.raise_for_status()
     return resp.json()

Nothing in these lines is wrong in isolation. The problem is what the diff does not show. The shared session, imported from a platform package, already has retries mounted on it:

platform/http.py · not in the difflayer 3
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

_retry = Retry(
    total=2,                              # 2 retries = 3 attempts
    backoff_factor=0.2,
    status_forcelist=[502, 503, 504],
    allowed_methods=None,                  # None = retry every verb, including POST
)
session = requests.Session()
session.mount("https://", HTTPAdapter(max_retries=_retry))

And the route in front of checkout, owned by the platform team, retries server errors too:

deploy/gateway/checkout-route.yaml · not in the difflayer 1
route:
  cluster: checkout
  timeout: 6s
  retry_policy:
    retry_on: "5xx,reset,connect-failure"
    num_retries: 2                         # 3 attempts
    per_try_timeout: 2s

Read together, the three files describe a much riskier change than the diff suggests:

IssueWhere it comes fromWhy it mattersWhat a reviewer should ask
27× amplificationGateway × decorator × clientA short inventory brownout becomes a sustained overloadWhich single layer should own retries for this path?
Retrying a non-idempotent POSTallowed_methods=None plus the decoratorA timed-out reservation may have succeeded; retrying can reserve stock twiceDoes /reservations accept an idempotency key?
Retrying client errorsraise_for_status() raises HTTPError, a RequestExceptionA 409 “out of stock” or 400 is retried three times and can never succeedWhich status codes are actually transient?
Synchronized retrieswait_fixed(0.5)Every failed caller retries at the same moment, in wavesIs the backoff exponential, capped, and jittered?
Retries past the deadlineInner retries can take ~10 s; gateway per-try timeout is 2 sWork continues for callers that already gave up, and overlaps the gateway’s own retryIs the remaining deadline checked before each attempt?

This is a textbook case of a change whose risk lives outside its lines. Assessing it means following the call path and its configuration, the same way you would assess the blast radius of a code change.

How do retries cause cascading failures?

A cascading failure is a failure that spreads because the system’s response to it creates more of the same failure. Retries are a direct path to that. They convert errors into load, and in an overload event, load is the cause of the errors.

THE RETRY STORM FEEDBACK LOOP 1 · Dependency slows down GC pause, hot shard, bad deploy, lost zone 2 · Calls time out or return 503 Error rate crosses what callers tolerate 3 · Every layer retries Attempts multiply: 3, 9, 27 per request 4 · Offered load exceeds capacity Queues grow, latency climbs, more timeouts Exit condition: offered load must fall below capacity. Unbounded retries push it the other way.
The trigger can be gone in seconds. The loop keeps the dependency overloaded because the callers generate the load.

Several things make the loop hard to escape once it starts:

The same shape shows up with caches: many callers acting at once on the same signal, turning a small event into a load spike on a shared dependency. See Cache Stampede: Why a Healthy Cache Can Overload Your Database for that variant.

A retry is a bet that the failure was random. During overload, the failure is not random, and every retry makes the bet worse.

What should never be retried?

Retries are safe only when two things are true: the failure is likely to be transient, and repeating the operation cannot cause a different outcome. Most retry storms and many data bugs come from ignoring one of those conditions.

ResponseTransient?Retry?Notes
Connection refused, DNS failure before sendingOftenYes, with backoffThe request never reached the server, so repeating it is safe even for writes
400, 401, 403, 404, 409, 422NoNoThe same request will fail the same way. Retrying only adds load
429 Too Many RequestsYes, but it is an overload signalOnly after Retry-After, within budgetThe server is explicitly asking for less traffic
502, 503, 504UsuallyAt one layer, with jittered backoff and budgetTreat a 503 that means “overloaded” as a signal to back off, not try again
500UnknownOnly for idempotent callsOften a deterministic bug; retries rarely help
Timeout after the request was sentUnknownOnly if idempotent or keyedThe server may have completed the work. The outcome is unknown, not failed
SHOULD THIS CALL RETRY? FOUR QUESTIONS, IN ORDER Q1 · ERROR TYPE Is the failure transient, not an overload signal? Q2 · OPERATION Is the call idempotent, or does it carry a key? Q3 · LAYER Is this the one layer that owns retries for the path? Q4 · BUDGET AND DEADLINE Budget available and time left before the deadline? Don’t retry Return the error upstream Add an idempotency key Then reconsider the retry Don’t retry here Pass the failure through Fail fast Shed load, keep capacity NONONONO YESYESYESYES Retry with capped exponential backoff and full jitter sleep = random(0, min(cap, base × 2^attempt)), then re-check the deadline
A retry is the last branch of the tree, not the default. Most failures during an overload should stop at Q1 or Q4.

Non-idempotent operations

A POST /reservations that timed out may have reserved the stock. A retry reserves it again. The same applies to charging a card, sending an email, or publishing a message. Before any write is retried, it needs an idempotency key that the server stores and checks, so a repeated request returns the original result instead of repeating the effect. The mechanics are the same as deduplicating webhooks; see How to Prevent Duplicate Webhook Processing.

Many HTTP client libraries retry only idempotent methods by default for this reason. In urllib3, the default allowed_methods set excludes POST; setting it to None, as the shared client above does, turns that protection off for every caller of that session.

Overload errors

Some errors mean “I am broken right now,” and others mean “I am too busy.” Retrying the first kind may help. Retrying the second kind is the retry storm. The SRE book’s chapter on Handling Overload describes a dedicated “overloaded; don’t retry” error that a backend returns once its own retries are exhausted, so the layers above do not multiply them. If your services cannot tell those cases apart, add the distinction: a specific status, header, or error code that callers must not retry.

Exponential backoff with jitter: why fixed delays synchronize retries

A fixed delay, like wait_fixed(0.5) in the PR above, means every caller that failed at the same moment retries at the same moment. The failures came in together, so the retries arrive together too, as waves that hit the dependency every half second. Exponential backoff spreads retries out over longer intervals. Jitter spreads them out across callers.

SAME RETRIES, DIFFERENT ARRIVAL PATTERN Fixed 500 ms no jitter Full jitter random(0, backoff) capacity capacity 0 s0.5 s1 s1.5 s2 s Same number of retries. Spread over time, they stay under capacity instead of arriving as waves.
Jitter does not reduce the number of retries. It removes the synchronization that turns them into spikes.

The AWS Architecture Blog post Exponential Backoff And Jitter compares several strategies with a simulation. The one most teams should default to is full jitter:

Capped exponential backoff with full jitterdefault choice
sleep = random_between(0, min(cap, base * 2 ** attempt))

# base = 100 ms, cap = 1 s:
# attempt 1: random in [0, 200 ms]
# attempt 2: random in [0, 400 ms]
# attempt 3: random in [0, 800 ms]
# attempt 4+: random in [0, 1 s]

In Python, tenacity provides this as wait_random_exponential(multiplier=..., max=...). gRPC’s built-in retry policy applies randomized exponential backoff through initialBackoff, maxBackoff, and backoffMultiplier. Whatever the library, check three things in review: the delay grows, it is capped, and it is randomized.

Backoff alone does not prevent a retry storm

Backoff and jitter shape when retries arrive. They do not limit how many there are. With 27 attempts per request, perfectly jittered retries are still 27 times the load. You need the limits in the next section too.

Retry budgets, deadlines, circuit breakers, and load shedding

Five controls, used together, keep retries from turning into a cascading failure. Each one answers a different question.

1. Retry at one layer only

Pick one layer per call path to own retries, usually the layer closest to the caller that knows the user’s deadline and whether the operation is safe to repeat. Every other layer passes failures through. This single rule changes the worst case from a product to a sum. If the gateway must retry, restrict it to failures where the request never reached the upstream (connection failures, refused streams), and disable retries in the libraries below.

2. Retry budgets

An attempt limit (“at most 3 tries”) bounds retries per request. A retry budget bounds retries per client or process: retries are allowed only while they stay under a fixed fraction of normal requests. The SRE book describes a per-request limit of three attempts combined with a per-client budget of about 10% of requests. When most requests are failing, the budget runs out quickly and the client stops retrying, which is the behavior you want during an overload. gRPC offers this as retryThrottling (maxTokens, tokenRatio) and Envoy offers a retry_budget on the cluster’s circuit breaker thresholds.

3. Deadlines and timeout propagation

Every request should carry a deadline, and every layer should check the remaining time before starting an attempt. A retry that cannot finish before the caller gives up is pure waste. gRPC propagates deadlines natively; over HTTP, pass the remaining budget in a header your services agree on, and derive each call’s timeout from it.

WITHOUT DEADLINES, RETRIES OUTLIVE THEIR CALLERS gateway per-try timeout: chain 1 abandoned Gateway checkout, chain 1 checkout, chain 2 attempt 1 (per-try 2 s) attempt 2 (retry) try 1 try 2 try 3, orphaned try 1 try 2 0 s1 s2 s3 s4 s Red: work for a caller that already gave up. It still uses inventory capacity, and it overlaps with chain 2, so inventory serves both chains for one user request.
A deadline checked before each attempt stops chain 1 at 2 seconds. Without it, abandoned retries add load at the worst possible time.

4. Circuit breakers

A circuit breaker watches the error rate for a dependency and, above a threshold, stops sending requests for a cool-down period, then lets a few probes through. A retry budget limits the extra load retries add; a circuit breaker removes most of the load entirely while the dependency is unhealthy, and fails fast so callers do not hold threads waiting. Resilience4j and Polly provide circuit breakers as library components; at the proxy layer, Envoy offers circuit breaking thresholds and outlier detection. Review that the breaker’s open state produces a clear, non-retryable error, or the layer above will retry the breaker.

5. Load shedding on the server

The callee protects itself as well. When a service is near capacity, it should reject excess requests cheaply and early, before doing expensive work, and return a response that tells callers not to retry immediately. Rejecting 20% of requests quickly keeps the other 80% fast. Accepting all of them makes all of them slow, which triggers timeouts and retries everywhere upstream.

The same change, done safely

Here is the checkout call rewritten so that retries live in exactly one place, with jittered backoff, a budget, a deadline, and an idempotency key. First, a small shared helper:

platform/retries.pyone layer, bounded
import random, threading, time

RETRYABLE_STATUS = {502, 503, 504}      # never 4xx; 429 needs Retry-After handling

class RetryBudget:
    """Token bucket: each request earns `ratio` tokens, each retry spends 1."""
    def __init__(self, ratio=0.1, max_tokens=10.0):
        self.ratio, self.max_tokens = ratio, max_tokens
        self.tokens = max_tokens
        self._lock = threading.Lock()

    def on_request(self):
        with self._lock:
            self.tokens = min(self.max_tokens, self.tokens + self.ratio)

    def try_spend(self):
        with self._lock:
            if self.tokens >= 1.0:
                self.tokens -= 1.0
                return True
            return False                     # budget empty: stop retrying

def call_with_retries(send, *, deadline, budget, max_attempts=3, base=0.1, cap=1.0):
    budget.on_request()
    attempt = 0
    while True:
        remaining = deadline - time.monotonic()
        if remaining <= 0:
            raise TimeoutError("deadline exceeded before attempt")
        resp = send(timeout=min(1.0, remaining))   # never outlive the caller
        attempt += 1
        if resp.status_code not in RETRYABLE_STATUS:
            return resp                           # success or a non-transient error
        if attempt >= max_attempts or not budget.try_spend():
            return resp                           # let the caller see the failure
        delay = random.uniform(0, min(cap, base * 2 ** attempt))   # full jitter
        if time.monotonic() + delay >= deadline:
            return resp
        time.sleep(delay)
checkout/inventory_client.pykeyed, deadline-aware
from platform.http import no_retry_session     # Retry(total=0): no hidden layer
from platform.retries import RetryBudget, call_with_retries

_inventory_budget = RetryBudget(ratio=0.1)

def reserve_stock(order_id, items, *, deadline):
    def send(timeout):
        return no_retry_session.post(
            f"{INVENTORY_URL}/reservations",
            json={"order_id": order_id, "items": items},
            headers={"Idempotency-Key": f"reserve-{order_id}"},  # safe to repeat
            timeout=timeout,
        )
    resp = call_with_retries(send, deadline=deadline, budget=_inventory_budget)
    resp.raise_for_status()
    return resp.json()

The sketch handles HTTP status codes only. A production version would also retry connection errors that happen before the request is sent, honor Retry-After on 429 and 503, and emit a metric for every retry and every budget refusal. Retries that are not measured are how a storm goes unnoticed until it is an outage.

And the gateway stops retrying server errors on this route, keeping only failures where the request never reached checkout, bounded by a budget:

deploy/gateway/checkout-route.yamlnarrowed + budgeted
route:
  cluster: checkout
  timeout: 3s
  retry_policy:
    retry_on: "connect-failure,refused-stream"   # request never reached checkout
    num_retries: 1
# in the checkout cluster definition:
circuit_breakers:
  thresholds:
    - priority: DEFAULT
      retry_budget:
        budget_percent: { value: 10.0 }
        min_retry_concurrency: 3

The worst case for one user request is now roughly 2 × 3 = 6 attempts on inventory, and far fewer in practice, because the budgets drain as soon as failures become widespread. More importantly, retrying a reservation can no longer reserve stock twice.

Retry storm checklist for pull request review

Use this whenever a pull request adds or changes a retry, a timeout, a client library, or gateway configuration. The first group is the one reviewers most often skip, because answering it requires looking outside the diff.

Layers on the call path
  • List every layer that retries this call: client SDK, gateway or mesh, service code, HTTP or gRPC client, driver
  • Multiply the attempt counts; flag anything above single digits
  • Name the one layer that should own retries
What gets retried
  • Only transient errors; no 4xx
  • No retry on explicit overload signals
  • Writes carry an idempotency key the server enforces
  • Library defaults for methods are not widened
How retries are spaced
  • Exponential, capped, and jittered backoff
  • No fixed-interval or immediate retries
  • Retry-After is honored when present
Limits and time
  • Small attempt limit plus a retry budget
  • Remaining deadline checked before each attempt
  • Per-attempt timeout shorter than the caller’s timeout
  • Circuit breaker errors are not retried upstream
  • Retries and budget refusals are emitted as metrics

A PR that fails the first group should be treated as a high-risk change regardless of its size. It fits the pattern in How to Identify High Risk Pull Requests: a small diff that changes behavior on a shared, critical path. For the wider review practice, see How to Review a Pull Request for Production Reliability Risks.

How Tomosu helps spot a retry storm before merge

A style-focused review of PR #4812 has little to object to. The decorator is idiomatic, the imports are tidy, the function is short. The risk only appears when the diff is read together with the shared client it calls through and the gateway route in front of it. That cross-layer reading is what Tomosu does on a pull request:

Illustrative pre-merge findingretry amplification
PR #4812 adds a retry layer to POST {INVENTORY_URL}/reservations
  layer 1  deploy/gateway/checkout-route.yaml  num_retries: 2      3 attempts
  layer 2  checkout/inventory_client.py        stop_after_attempt  3 attempts  (new)
  layer 3  platform/http.py                    Retry(total=2)      3 attempts
  worst case per user request: 27 calls to inventory-service
  also: POST retried without Idempotency-Key; 4xx retried; wait_fixed (no jitter)
  evidence to check: inventory 503 rate during deploys, retry metrics, deadline

These findings feed the Fragility Index and the overall Production Reliability Index, so a three-line retry decorator is weighed by what it does to the call path, not by its line count. It fits into the broader practice of pre merge reliability analysis.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

What is a retry storm?

A retry storm is a feedback loop in which clients retry failed requests to a service that is failing because it is overloaded. The retries add more load, which causes more failures, which cause more retries. It often starts from a brief slowdown and continues after the original trigger is gone, because the callers now generate the excess traffic.

How do retries cause cascading failures in microservices?

When several layers each retry, attempts multiply: three layers with three attempts each can send 27 requests to the bottom service for one user action. During an overload nearly every attempt fails, so every layer uses its full allowance. Callers waiting on retries also hold threads and connections, lose capacity themselves, and get retried by their own callers, so the overload spreads upward.

How many retries is too many?

Count attempts across the whole call path, not per layer. A common guideline is a small number of attempts, such as three, at a single layer, combined with a retry budget that keeps retries to roughly 10% of normal requests. If multiplying the attempt counts of every retrying layer on a path gives more than single digits, the path needs to be simplified.

What is exponential backoff with jitter?

Exponential backoff increases the delay between retries, typically doubling it each attempt up to a cap. Jitter randomizes each delay so that callers that failed together do not retry together. With full jitter, the delay is a random value between zero and min(cap, base × 2^attempt). Backoff and jitter control when retries happen, not how many, so they should be combined with attempt limits and a retry budget.

Should I retry on HTTP 429 or 503?

Treat both as possible overload signals. For 429, wait at least as long as the Retry-After header says, and only retry within your retry budget. For 503, retry at one layer with jittered backoff if the error is transient, but do not retry if the service indicates it is overloaded. Never retry 4xx errors such as 400, 404, or 409, because the same request will fail the same way.

Is it safe to retry a POST request?

Only if the operation is idempotent or carries an idempotency key that the server stores and checks. A POST that timed out may have succeeded on the server, so a blind retry can create a duplicate order, reservation, or charge. It is safe to retry a POST when the connection failed before the request was sent, because the server never received it.

What is the difference between a retry budget and a circuit breaker?

A retry budget limits how many retries a client may send, usually as a fraction of its normal requests, so retries cannot multiply load during an outage. A circuit breaker stops sending requests to an unhealthy dependency for a period and fails fast instead. The budget caps the extra load from retries; the breaker removes most of the load entirely while the dependency recovers. They work best together.

Where should retries live in a microservice call chain?

At one layer per call path, usually the one closest to the caller that knows the request deadline and whether the operation is safe to repeat. Other layers should pass failures through instead of retrying. If a gateway or service mesh retries at all, limit it to failures where the request never reached the upstream, and disable retries in client libraries below the owning layer.


A retry storm is written one reasonable line at a time, in different files, by different teams. Tomosu reads those layers together, so the multiplication is visible before merge instead of during the incident. Assess your repository →