Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · API resilience

Circuit Breaker vs. Retry vs. Load Shedding: Which Failure Needs Which Control?

Tomosu AI·14 min read·

Retries, circuit breakers, and load shedding are often listed together as “resilience patterns”, as if any of them would do. They answer different failures. Applied to the wrong one, each makes the incident worse: a retry deepens an overload, a breaker hides a bug, and shedding in the wrong place drops the traffic you most needed to serve.

Quick answer

Choose by the failure, not the pattern. A retry is for a brief, random failure on an idempotent call. A circuit breaker is for a dependency that is failing or slow for many calls in a row: the caller stops calling it and fails fast. Load shedding is for a server receiving more work than it can do: it rejects the excess early so the rest stays fast.

This guide explains what each control does, where it lives, which failure it fits, and which failure it makes worse. It then shows how to combine them in one call path, with Resilience4j configuration as the worked example. The retry side is covered in depth in Retry Storm in Microservices: How to Spot One Before Merge; this post is about choosing between the controls.

What do retries, circuit breakers, and load shedding each do?

Each control protects something different, and that is the fastest way to tell them apart.

WHERE EACH CONTROL LIVES CALLER DEPENDENCY (CALLEE) Retry rides out a brief, random failure Circuit breaker stops calling a failing dependency Timeout bounds how long one attempt waits Load shedding rejects excess work early and cheaply Concurrency limit caps work in progress and queue length Handler does the real work, within its deadline request Retries and breakers protect the caller’s request. Shedding protects the server from its callers.
The request leaves the caller through its timeout, and the first thing it meets at the dependency is the load shedder.
ControlLives inTriggered byProtectsMakes worse if misused
RetryCaller, per callOne failed attemptA single requestOverload: every retry is extra load on a service that is already failing from load
Circuit breakerCaller, per dependencyFailure or slow-call rate over a windowThe caller’s threads, connections, and latencyBugs and bad requests: counting 4xx as failures opens the breaker on healthy dependencies
Load sheddingCallee, at the entry pointThe server’s own load: in-flight work, queue time, CPUThe server and the requests it acceptsMisplaced shedding: rejecting after expensive work, or dropping critical traffic first
Timeout (the base)BothElapsed timeBounded waitingInverted or missing timeouts: every other control depends on them

Two neighbors often get confused with these. A rate limiter enforces a quota per client whether or not the server is busy; load shedding responds to the server’s actual load. A bulkhead reserves a separate pool of threads or connections per dependency, so one slow dependency cannot use them all. Both are useful, and neither replaces the three above.

Which failure needs which control?

Classify the failure by two questions: whose problem is it (the request, a dependency, or your own service), and is it brief or sustained? The answer picks the control.

WHICH CONTROL? FOUR QUESTIONS, IN ORDER Q1 · THE REQUEST Is the error caused by the request itself (4xx)? Q2 · YOUR OWN SERVICE Is your service receiving more than it can serve? Q3 · A DEPENDENCY, SUSTAINED Is one dependency failing or slow for many calls? Q4 · A DEPENDENCY, BRIEF Is the failure brief and random, the call idempotent? No resilience control Fix the caller; don’t retry Load shedding Reject early: 503 or 429 Circuit breaker Fail fast, serve a fallback Retry, once or twice Jittered, within the budget YESYESYESYES NONONONO Fail fast and surface the error Unknown or persistent failures: don’t retry blindly. Return a clear error, alert, and investigate.
Ask about your own load before asking about the dependency. A retry is the last branch, not the default.
FailureRight controlWrong control, and why
Connection reset during a pod restartRetry once, with jitterCircuit breaker: one blip should not stop all traffic
Dependency returns 5xx for most callsCircuit breaker with a fallbackRetry: multiplies load on a service that is already down
Dependency is slow, not erroringTimeout, plus a breaker with a slow-call threshold and a bulkheadLonger timeouts: threads pile up waiting
Your service gets more traffic than it can serveLoad shedding, then scale outUnbounded queues: every request waits and then times out
One bad instance behind a load balancerOutlier detection or health-based ejection of that instanceService-wide breaker: cuts off the healthy instances too
Third-party API returns 429Honor Retry-After; client-side rate limitImmediate retry: asks again at exactly the wrong moment
400, 404, 409, 422None; return the errorAny of them: the same request fails the same way

When is a retry the right control?

A retry is right when three things are true: the failure is likely to be random and brief, repeating the call cannot cause a different outcome, and there is time left in the caller’s deadline. Connection resets during a deploy, a dropped packet, and a request that reached a draining instance are the classic cases. Most of them succeed on the first retry.

The limits matter more than the retry. Keep attempts small (one retry is often enough), use capped exponential backoff with jitter, cap retries per client with a budget, retry at one layer only, and never retry non-idempotent writes without an idempotency key. The retry storm post covers each of those in detail, including why nested retries multiply.

The retry that looks like resilience

A retry cannot tell “random failure” from “overloaded”. During overload nearly every attempt fails, so every caller uses its full retry allowance at the moment the dependency has the least capacity. A retry without a budget and a breaker is a load multiplier waiting for a bad day.

When should you use a circuit breaker?

Use a circuit breaker when a dependency can fail or slow down for a sustained period, and your service has something better to do than wait: a fallback, a cached value, a degraded response, or a fast, clear error. The breaker turns “every request waits for the timeout and then fails” into “requests fail in microseconds while the dependency recovers”. Martin Fowler’s CircuitBreaker article describes the original pattern.

CIRCUIT BREAKER STATES CLOSED Calls pass through; outcomes are recorded OPEN Calls fail fast; the fallback is served HALF-OPEN A few probe calls are let through failure or slow-call rate ≥ threshold over the sliding window after wait duration probes succeed probes fail
In Resilience4j terms: failureRateThreshold and slowCallRateThreshold open it, waitDurationInOpenState holds it open, and permittedNumberOfCallsInHalfOpenState sets the number of probes.

Four settings decide whether a breaker helps or hurts:

A breaker without a fallback is still useful, since failing in microseconds beats failing after a timeout. But its real value comes from what the caller does instead: serve cached data, hide an optional widget, queue the work for later, or return a clear error the layer above will not retry. The Resilience4j CircuitBreaker documentation lists every setting and its default.

When do you need load shedding?

You need load shedding whenever a service can receive more traffic than it can serve, which is every service exposed to real users, retries, or batch jobs. Without it, excess requests queue. Queued requests wait, time out at the caller, and get retried, while the server keeps working on requests nobody is still waiting for. Throughput of useful responses falls just when demand peaks.

GOODPUT UNDER OVERLOAD, WITH AND WITHOUT SHEDDING capacity goodput (successful responses/s) With load shedding goodput holds; excess gets a fast 503 Without shedding queues grow, timeouts and retries add load 50%100%150%200% offered load, as a share of capacity →
Shedding does not add capacity. It stops the server from spending its capacity on requests that will time out anyway.

Good load shedding has four properties:

The most common missing shedder is an unbounded queue. A fixed thread pool in Java with the default factory method accepts work forever:

OrderServer.javano shedding: the queue never says no
// newFixedThreadPool uses an unbounded LinkedBlockingQueue
ExecutorService workers = Executors.newFixedThreadPool(50);

// Under overload, work waits here until the caller has already timed out
workers.execute(() -> handle(request));
OrderServer.javabounded, rejects cheaply
ThreadPoolExecutor workers = new ThreadPoolExecutor(
        50, 50, 0L, TimeUnit.MILLISECONDS,
        new ArrayBlockingQueue<>(100),           // at most 100 requests waiting
        new ThreadPoolExecutor.AbortPolicy());   // full: throw instead of queueing

try {
    workers.execute(() -> handle(request));
} catch (RejectedExecutionException overloaded) {
    reply(request, 503, Map.of("Retry-After", "1"));   // shed before doing any work
}

A bounded queue is the simplest form of shedding. Adaptive concurrency limits, which adjust the limit from observed latency, and queue-time checks, which drop requests that have already waited longer than their caller will, are refinements of the same idea. The Google SRE book chapter on handling overload and the Amazon Builders’ Library article on using load shedding to avoid overload both go deeper.

How do you combine them in one call path?

Real call paths need all three, each doing its own job. Put them in this order, from the outside of a call inward:

  1. Deadline. The overall time the request has left, propagated from the caller.
  2. Retry. Few attempts, with a budget, only for retryable errors, and only while the deadline allows.
  3. Circuit breaker. Checked on every attempt, so an open breaker fails the attempt instantly. The retry must not retry a breaker rejection.
  4. Bulkhead or concurrency limit. Caps how many calls to this dependency can be in flight.
  5. Per-attempt timeout. Bounds each individual call, derived from the remaining deadline.

And on the server side, load shedding sits in front of all of the service’s own work. Resilience4j’s Spring Boot integration applies its annotations in a similar default order, with Retry outermost: Retry, then CircuitBreaker, RateLimiter, TimeLimiter, and Bulkhead. The companion post How to Set Timeouts Across a Chain of Microservices covers how the deadline and per-attempt timeouts fit together.

Here is a common configuration that gets the combination wrong, followed by one that gets it right.

application.ymlretries everything, no breaker
resilience4j:
  retry:
    instances:
      inventory:
        max-attempts: 5        # 4 retries, on every exception, including 4xx
        wait-duration: 1s      # fixed delay: callers retry in waves
application.ymlbreaker + narrow, jittered retry
resilience4j:
  circuitbreaker:
    instances:
      inventory:
        sliding-window-type: COUNT_BASED
        sliding-window-size: 50
        minimum-number-of-calls: 20
        failure-rate-threshold: 50
        slow-call-duration-threshold: 2s
        slow-call-rate-threshold: 80
        wait-duration-in-open-state: 20s
        permitted-number-of-calls-in-half-open-state: 5
        ignore-exceptions:
          # a business answer, not a failure
          - com.example.inventory.OutOfStockException
  retry:
    instances:
      inventory:
        max-attempts: 2                # one retry
        wait-duration: 200ms
        enable-randomized-wait: true   # jitter
        randomized-wait-factor: 0.5
        retry-exceptions:
          - java.net.ConnectException   # only what is plausibly transient
        ignore-exceptions:
          - io.github.resilience4j.circuitbreaker.CallNotPermittedException
StockService.java
@CircuitBreaker(name = "inventory", fallbackMethod = "stockUnknown")
@Retry(name = "inventory")
public Stock getStock(String sku) {
    // inventoryClient has its own connect and read timeouts
    return inventoryClient.getStock(sku);
}

// Breaker open: show "availability unknown" instead of failing the page
private Stock stockUnknown(String sku, CallNotPermittedException e) {
    return Stock.unknown(sku);
}

The numbers are illustrative; take yours from the dependency’s measured error and latency profile. What matters is the shape: the breaker judges the dependency’s health, the retry handles only plausibly transient errors and does not fight the breaker, business answers do not count as failures, and the open state has a fallback.

A retry bets the failure is random. A breaker bets it is not. Load shedding bets the problem is you.

What goes wrong when the control doesn’t match the failure?

Most resilience incidents are not missing controls. They are controls aimed at the wrong failure.

Each of these is visible in code and configuration before it is visible in production, but rarely in a single diff. The retry lives in one file, the breaker config in another, the HTTP client’s defaults in a shared library, and the server’s queue in a framework setting. Reviewing them together is the point of How to Review a Pull Request for Production Reliability Risks. Developers ask this question often; a Stack Overflow search for circuit breaker vs retry shows how often the controls get treated as interchangeable. When you look at how far a misapplied control can spread, How to Assess the Blast Radius of a Code Change is the companion method.

How Tomosu helps

Tomosu analyzes a repository and each pull request for production reliability risk. For resilience controls, the useful question is not whether a pattern exists but whether it fits the failure on that path:

These findings feed the Production Reliability Index. To run it on your own code, follow the repository scan guide.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

What is the difference between a circuit breaker and a retry?

A retry repeats one failed call in the hope that the failure was random and brief. A circuit breaker watches the failure rate of many calls to a dependency and, when it crosses a threshold, stops calling that dependency for a while and fails fast. Retries help with isolated blips; circuit breakers help when a dependency is failing or slow for a sustained period.

When should I use load shedding instead of a circuit breaker?

Use load shedding when your own service is receiving more work than it can serve; it runs in the server and rejects excess requests early. Use a circuit breaker when a dependency you call is failing or slow; it runs in the caller and stops calls to that dependency. Many systems need both, one on each side of a call.

Should I use a circuit breaker and retries together?

Yes, if they are ordered and scoped correctly. The retry should wrap the circuit breaker so each attempt checks the breaker, it should retry only plausibly transient errors, and it must not retry the breaker’s own rejection. Resilience4j’s Spring Boot integration applies Retry outside CircuitBreaker by default.

Why do retries make overload worse?

During overload nearly every attempt fails, so every caller uses its full retry allowance at the moment the dependency has the least spare capacity. Retries convert errors into extra load, and in an overload the load is the cause of the errors. Budgets, backoff with jitter, and circuit breakers limit that effect.

Should a circuit breaker count 4xx errors as failures?

Usually not. Client errors such as 400, 404, 409, and 422, and business answers such as out of stock, mean the dependency is working correctly. Counting them can open the breaker on a healthy dependency because one caller is sending bad requests. Count timeouts, connection errors, and 5xx responses, and consider slow calls too.

What should a service return when it sheds load?

Return a cheap, explicit rejection before doing expensive work: typically HTTP 503 with a Retry-After header for server overload, or 429 when a specific client exceeds its quota. The response should tell well-behaved callers to back off rather than retry immediately.

Is rate limiting the same as load shedding?

No. Rate limiting enforces a fixed quota per client or key, whether or not the server is busy. Load shedding reacts to the server’s actual load, such as requests in flight or time spent queued, and rejects work only when capacity runs out. Rate limits provide fairness between clients; shedding protects the server from overload.


A resilience control is only as good as its match to the failure. Tomosu reads the retries, breakers, and queues across your repository and shows where they don’t fit. Assess your repository →