Retries, circuit breakers, and load shedding are often listed together as “resilience patterns”, as if any of them would do. They answer different failures. Applied to the wrong one, each makes the incident worse: a retry deepens an overload, a breaker hides a bug, and shedding in the wrong place drops the traffic you most needed to serve.
Choose by the failure, not the pattern. A retry is for a brief, random failure on an idempotent call. A circuit breaker is for a dependency that is failing or slow for many calls in a row: the caller stops calling it and fails fast. Load shedding is for a server receiving more work than it can do: it rejects the excess early so the rest stays fast.
- Transient blip: retry once or twice, jittered, within the deadline.
- Failing or slow dependency: circuit breaker with a fallback, plus timeouts.
- Overloaded server: load shedding with a fast 503 or 429.
- Bad request (4xx): none of the three. Fix the caller.
This guide explains what each control does, where it lives, which failure it fits, and which failure it makes worse. It then shows how to combine them in one call path, with Resilience4j configuration as the worked example. The retry side is covered in depth in Retry Storm in Microservices: How to Spot One Before Merge; this post is about choosing between the controls.
What do retries, circuit breakers, and load shedding each do?
Each control protects something different, and that is the fastest way to tell them apart.
- A retry repeats a failed call in the hope that the failure was random. It protects one request from a brief fault. It runs in the caller.
- A circuit breaker tracks the recent failure rate of calls to one dependency, and when it crosses a threshold, stops making those calls for a while and fails immediately. It protects the caller from waiting on a broken dependency, and gives the dependency room to recover. It runs in the caller, one per dependency.
- Load shedding rejects some incoming requests when a server is at capacity, before doing any real work on them. It protects the server from its callers, so the requests it does accept stay fast. It runs in the callee.
| Control | Lives in | Triggered by | Protects | Makes worse if misused |
|---|---|---|---|---|
| Retry | Caller, per call | One failed attempt | A single request | Overload: every retry is extra load on a service that is already failing from load |
| Circuit breaker | Caller, per dependency | Failure or slow-call rate over a window | The caller’s threads, connections, and latency | Bugs and bad requests: counting 4xx as failures opens the breaker on healthy dependencies |
| Load shedding | Callee, at the entry point | The server’s own load: in-flight work, queue time, CPU | The server and the requests it accepts | Misplaced shedding: rejecting after expensive work, or dropping critical traffic first |
| Timeout (the base) | Both | Elapsed time | Bounded waiting | Inverted or missing timeouts: every other control depends on them |
Two neighbors often get confused with these. A rate limiter enforces a quota per client whether or not the server is busy; load shedding responds to the server’s actual load. A bulkhead reserves a separate pool of threads or connections per dependency, so one slow dependency cannot use them all. Both are useful, and neither replaces the three above.
Which failure needs which control?
Classify the failure by two questions: whose problem is it (the request, a dependency, or your own service), and is it brief or sustained? The answer picks the control.
| Failure | Right control | Wrong control, and why |
|---|---|---|
| Connection reset during a pod restart | Retry once, with jitter | Circuit breaker: one blip should not stop all traffic |
| Dependency returns 5xx for most calls | Circuit breaker with a fallback | Retry: multiplies load on a service that is already down |
| Dependency is slow, not erroring | Timeout, plus a breaker with a slow-call threshold and a bulkhead | Longer timeouts: threads pile up waiting |
| Your service gets more traffic than it can serve | Load shedding, then scale out | Unbounded queues: every request waits and then times out |
| One bad instance behind a load balancer | Outlier detection or health-based ejection of that instance | Service-wide breaker: cuts off the healthy instances too |
| Third-party API returns 429 | Honor Retry-After; client-side rate limit | Immediate retry: asks again at exactly the wrong moment |
| 400, 404, 409, 422 | None; return the error | Any of them: the same request fails the same way |
When is a retry the right control?
A retry is right when three things are true: the failure is likely to be random and brief, repeating the call cannot cause a different outcome, and there is time left in the caller’s deadline. Connection resets during a deploy, a dropped packet, and a request that reached a draining instance are the classic cases. Most of them succeed on the first retry.
The limits matter more than the retry. Keep attempts small (one retry is often enough), use capped exponential backoff with jitter, cap retries per client with a budget, retry at one layer only, and never retry non-idempotent writes without an idempotency key. The retry storm post covers each of those in detail, including why nested retries multiply.
A retry cannot tell “random failure” from “overloaded”. During overload nearly every attempt fails, so every caller uses its full retry allowance at the moment the dependency has the least capacity. A retry without a budget and a breaker is a load multiplier waiting for a bad day.
When should you use a circuit breaker?
Use a circuit breaker when a dependency can fail or slow down for a sustained period, and your service has something better to do than wait: a fallback, a cached value, a degraded response, or a fast, clear error. The breaker turns “every request waits for the timeout and then fails” into “requests fail in microseconds while the dependency recovers”. Martin Fowler’s CircuitBreaker article describes the original pattern.
Four settings decide whether a breaker helps or hurts:
- What counts as a failure. Timeouts, connection errors, and 5xx should count. Client errors and business answers (“out of stock”, “not found”) should not; in Resilience4j that is
ignoreExceptions. Otherwise a spike of bad requests opens the breaker on a healthy dependency. - Slow calls. A dependency that stops erroring but takes 10 seconds per call is just as harmful. Resilience4j’s
slowCallDurationThresholdandslowCallRateThresholdlet slowness open the breaker too. - Minimum calls. A threshold on a tiny sample flaps.
minimumNumberOfCallsstops the breaker from opening on two failures out of three calls. - Scope. One breaker per dependency (sometimes per endpoint), not one per service. For one bad instance among many, instance-level ejection such as Envoy’s outlier detection fits better than a breaker that cuts off the whole cluster.
A breaker without a fallback is still useful, since failing in microseconds beats failing after a timeout. But its real value comes from what the caller does instead: serve cached data, hide an optional widget, queue the work for later, or return a clear error the layer above will not retry. The Resilience4j CircuitBreaker documentation lists every setting and its default.
When do you need load shedding?
You need load shedding whenever a service can receive more traffic than it can serve, which is every service exposed to real users, retries, or batch jobs. Without it, excess requests queue. Queued requests wait, time out at the caller, and get retried, while the server keeps working on requests nobody is still waiting for. Throughput of useful responses falls just when demand peaks.
Good load shedding has four properties:
- It rejects early. The check happens at the entry point, before parsing large bodies, opening transactions, or calling dependencies. A rejection should cost almost nothing.
- It uses the server’s own signal. Requests in flight, time spent queued, or CPU, rather than a fixed request rate chosen months ago.
- It tells callers not to hammer it. A 503 with
Retry-After, or a 429 for per-client limits, so well-behaved callers back off instead of retrying immediately. - It sheds the least important work first. Background refreshes, prefetches, and batch traffic go before checkout and login, and load balancer health checks are exempt so a busy instance is not mistaken for a dead one. Priority is what makes shedding a product decision, not only an infrastructure one.
The most common missing shedder is an unbounded queue. A fixed thread pool in Java with the default factory method accepts work forever:
// newFixedThreadPool uses an unbounded LinkedBlockingQueue
ExecutorService workers = Executors.newFixedThreadPool(50);
// Under overload, work waits here until the caller has already timed out
workers.execute(() -> handle(request));
ThreadPoolExecutor workers = new ThreadPoolExecutor(
50, 50, 0L, TimeUnit.MILLISECONDS,
new ArrayBlockingQueue<>(100), // at most 100 requests waiting
new ThreadPoolExecutor.AbortPolicy()); // full: throw instead of queueing
try {
workers.execute(() -> handle(request));
} catch (RejectedExecutionException overloaded) {
reply(request, 503, Map.of("Retry-After", "1")); // shed before doing any work
}
A bounded queue is the simplest form of shedding. Adaptive concurrency limits, which adjust the limit from observed latency, and queue-time checks, which drop requests that have already waited longer than their caller will, are refinements of the same idea. The Google SRE book chapter on handling overload and the Amazon Builders’ Library article on using load shedding to avoid overload both go deeper.
How do you combine them in one call path?
Real call paths need all three, each doing its own job. Put them in this order, from the outside of a call inward:
- Deadline. The overall time the request has left, propagated from the caller.
- Retry. Few attempts, with a budget, only for retryable errors, and only while the deadline allows.
- Circuit breaker. Checked on every attempt, so an open breaker fails the attempt instantly. The retry must not retry a breaker rejection.
- Bulkhead or concurrency limit. Caps how many calls to this dependency can be in flight.
- Per-attempt timeout. Bounds each individual call, derived from the remaining deadline.
And on the server side, load shedding sits in front of all of the service’s own work. Resilience4j’s Spring Boot integration applies its annotations in a similar default order, with Retry outermost: Retry, then CircuitBreaker, RateLimiter, TimeLimiter, and Bulkhead. The companion post How to Set Timeouts Across a Chain of Microservices covers how the deadline and per-attempt timeouts fit together.
Here is a common configuration that gets the combination wrong, followed by one that gets it right.
resilience4j:
retry:
instances:
inventory:
max-attempts: 5 # 4 retries, on every exception, including 4xx
wait-duration: 1s # fixed delay: callers retry in waves
resilience4j:
circuitbreaker:
instances:
inventory:
sliding-window-type: COUNT_BASED
sliding-window-size: 50
minimum-number-of-calls: 20
failure-rate-threshold: 50
slow-call-duration-threshold: 2s
slow-call-rate-threshold: 80
wait-duration-in-open-state: 20s
permitted-number-of-calls-in-half-open-state: 5
ignore-exceptions:
# a business answer, not a failure
- com.example.inventory.OutOfStockException
retry:
instances:
inventory:
max-attempts: 2 # one retry
wait-duration: 200ms
enable-randomized-wait: true # jitter
randomized-wait-factor: 0.5
retry-exceptions:
- java.net.ConnectException # only what is plausibly transient
ignore-exceptions:
- io.github.resilience4j.circuitbreaker.CallNotPermittedException
@CircuitBreaker(name = "inventory", fallbackMethod = "stockUnknown")
@Retry(name = "inventory")
public Stock getStock(String sku) {
// inventoryClient has its own connect and read timeouts
return inventoryClient.getStock(sku);
}
// Breaker open: show "availability unknown" instead of failing the page
private Stock stockUnknown(String sku, CallNotPermittedException e) {
return Stock.unknown(sku);
}
The numbers are illustrative; take yours from the dependency’s measured error and latency profile. What matters is the shape: the breaker judges the dependency’s health, the retry handles only plausibly transient errors and does not fight the breaker, business answers do not count as failures, and the open state has a fallback.
A retry bets the failure is random. A breaker bets it is not. Load shedding bets the problem is you.
What goes wrong when the control doesn’t match the failure?
Most resilience incidents are not missing controls. They are controls aimed at the wrong failure.
- Retries against overload. The dependency is slow because it is saturated, and every caller multiplies its traffic. This is the retry storm.
- A breaker that counts client errors. A bad deploy of a caller sends malformed requests, the dependency correctly answers 400, and the breaker opens for every other caller too.
- A breaker with no slow-call rule. The dependency stops erroring and starts taking 20 seconds. The breaker stays closed while the caller’s threads fill up.
- Retrying the breaker. The retry wraps the breaker and treats “call not permitted” as retryable, so an open breaker becomes a loop of instant failures.
- Shedding after the expensive part. The server rejects requests only after authenticating, parsing, and querying. Rejections cost almost as much as successes, so shedding saves nothing.
- Shedding the wrong traffic. A limit applied equally to everything rejects checkout as readily as background polling.
Each of these is visible in code and configuration before it is visible in production, but rarely in a single diff. The retry lives in one file, the breaker config in another, the HTTP client’s defaults in a shared library, and the server’s queue in a framework setting. Reviewing them together is the point of How to Review a Pull Request for Production Reliability Risks. Developers ask this question often; a Stack Overflow search for circuit breaker vs retry shows how often the controls get treated as interchangeable. When you look at how far a misapplied control can spread, How to Assess the Blast Radius of a Code Change is the companion method.
How Tomosu helps
Tomosu analyzes a repository and each pull request for production reliability risk. For resilience controls, the useful question is not whether a pattern exists but whether it fits the failure on that path:
- Retry configuration: which exceptions and status codes are retried, attempt counts, backoff, and whether several layers retry the same call.
- Calls to dependencies without a breaker or fallback on request paths where a failing dependency would hold threads or connections.
- Unbounded queues and pools on the server side, where overload would queue instead of shedding.
- Blast radius: which endpoints share a dependency or a pool, so a missing control on a hot path is weighed accordingly.
These findings feed the Production Reliability Index. To run it on your own code, follow the repository scan guide.
Scan your repository with Tomosu →
Key takeaways
- Choose the control by the failure: brief and random, sustained in a dependency, or overload in your own service.
- Retries protect one request from a transient fault. Keep them few, jittered, budgeted, and idempotent.
- Circuit breakers protect the caller from a failing or slow dependency. Count timeouts and 5xx, not client errors, and include slow calls.
- Load shedding protects the server from its callers. Reject early and cheaply, signal
Retry-After, and shed low-priority work first. - Combine them in order: deadline, retry, breaker, bulkhead, per-attempt timeout. Never retry a breaker rejection.
- Client errors need none of the three. Fix the caller.
Frequently asked questions
What is the difference between a circuit breaker and a retry?
A retry repeats one failed call in the hope that the failure was random and brief. A circuit breaker watches the failure rate of many calls to a dependency and, when it crosses a threshold, stops calling that dependency for a while and fails fast. Retries help with isolated blips; circuit breakers help when a dependency is failing or slow for a sustained period.
When should I use load shedding instead of a circuit breaker?
Use load shedding when your own service is receiving more work than it can serve; it runs in the server and rejects excess requests early. Use a circuit breaker when a dependency you call is failing or slow; it runs in the caller and stops calls to that dependency. Many systems need both, one on each side of a call.
Should I use a circuit breaker and retries together?
Yes, if they are ordered and scoped correctly. The retry should wrap the circuit breaker so each attempt checks the breaker, it should retry only plausibly transient errors, and it must not retry the breaker’s own rejection. Resilience4j’s Spring Boot integration applies Retry outside CircuitBreaker by default.
Why do retries make overload worse?
During overload nearly every attempt fails, so every caller uses its full retry allowance at the moment the dependency has the least spare capacity. Retries convert errors into extra load, and in an overload the load is the cause of the errors. Budgets, backoff with jitter, and circuit breakers limit that effect.
Should a circuit breaker count 4xx errors as failures?
Usually not. Client errors such as 400, 404, 409, and 422, and business answers such as out of stock, mean the dependency is working correctly. Counting them can open the breaker on a healthy dependency because one caller is sending bad requests. Count timeouts, connection errors, and 5xx responses, and consider slow calls too.
What should a service return when it sheds load?
Return a cheap, explicit rejection before doing expensive work: typically HTTP 503 with a Retry-After header for server overload, or 429 when a specific client exceeds its quota. The response should tell well-behaved callers to back off rather than retry immediately.
Is rate limiting the same as load shedding?
No. Rate limiting enforces a fixed quota per client or key, whether or not the server is busy. Load shedding reacts to the server’s actual load, such as requests in flight or time spent queued, and rejects work only when capacity runs out. Rate limits provide fairness between clients; shedding protects the server from overload.
A resilience control is only as good as its match to the failure. Tomosu reads the retries, breakers, and queues across your repository and shows where they don’t fit. Assess your repository →