Every service in the chain has a timeout, and each one was chosen by a different team at a different time. The gateway gives up after 10 seconds, the service behind it waits 30, and the one behind that waits forever. When a dependency slows down, those numbers decide whether the slowdown stays local or spreads.
Set timeouts across a chain of microservices from the edge inward: each hop must give up before its caller does. Start from the user-facing budget, subtract each hop’s own work and a margin, and pass the remaining time downstream as a deadline. Every callee checks that deadline before it starts work and stops when it expires.
- No infinite defaults: every outbound call gets a connect timeout and an overall timeout.
- Inner shorter than outer: a callee’s timeout, retries included, fits inside its caller’s.
- Measured, not guessed: per-try timeouts come from the dependency’s healthy tail latency.
- Propagate the deadline: with gRPC deadlines or a header, and cancel abandoned work.
The previous post in this cluster, API Timeouts After Deployment: A Practical Triage Checklist, is about restoring service once timeouts are firing. This one is about the design that keeps a slow dependency from turning into a slow system: how to choose the numbers, how to make them agree across services, and how to check them in review.
Why do timeouts across a microservice chain go wrong?
They go wrong because each timeout is set locally, but the failure is global. A timeout limits how long one caller waits for one callee. In a chain, the timeouts only work if they are nested: each hop has to finish, or give up, before the hop above it stops waiting. When they are not nested, the chain has inverted timeouts, and the inner services keep working long after the outer ones have returned an error.
Inverted timeouts cause three problems at once:
- Orphaned work. Inner services spend CPU, threads, and database connections on requests whose callers already returned an error. During a slowdown, that is exactly the capacity the system is short of.
- Misleading errors. The edge reports a 504, the inner services report nothing wrong, and the traces show successful calls that nobody used. Triage starts in the wrong place.
- Amplified retries. If the edge or the client retries, a new chain starts while the old one is still running. The inner services now serve two chains for one user. Retry Storm in Microservices: How to Spot One Before Merge shows how quickly that multiplies.
The opposite mistake is just as common: timeouts so short that a healthy dependency’s normal tail latency trips them, producing errors and retries on a good day. The goal is not “short” or “long”. It is nested and measured.
What is a timeout budget?
A timeout budget is the total time a request is allowed to take, divided among the hops and attempts that serve it. It starts at the edge, where the product decides how long a user or client will wait, and shrinks at every hop by the time that hop has already spent.
For any hop, the rule is simple arithmetic: the time you allow a downstream call, multiplied by the attempts you allow, plus backoff, plus any other calls and your own work, plus a margin to send the response back, must be no more than the time your caller gave you.
Two consequences fall out of the arithmetic. First, retries are not free: two attempts at 2 seconds cost 4 seconds of the caller’s budget, so adding a retry without shortening the per-try timeout silently breaks the nesting. Second, the budget runs out faster the deeper the chain. A five-hop chain with a 2-second edge budget leaves very little for the last hop, which is a design signal in itself: the deepest calls need to be the fastest, cached, or moved off the request path.
Which timeouts does each call actually need?
“The timeout” is usually several settings. A call can hang while connecting, during a TLS handshake, while waiting for response headers, or while reading a slow body. A read timeout alone does not bound the total: a server that sends one byte every few seconds can keep a call alive indefinitely. Each outbound call needs at least:
- A connect timeout, short (hundreds of milliseconds inside a data center), so an unreachable host fails fast.
- An overall timeout or deadline covering the whole call, including reading the body.
- A database statement timeout for queries, so a slow query cannot hold a connection past the request’s life.
The defaults are the trap. Many common clients do not time out unless you tell them to:
| Client or layer | Default | What to set |
|---|---|---|
Python requests | No timeout unless you pass timeout | timeout=(connect, read) on every call, or a session wrapper that enforces it |
Go net/http | http.Client with zero Timeout never times out; http.DefaultClient is one | A context deadline per request, plus dial, TLS, and response header timeouts on the transport |
Java java.net.http.HttpClient | No connect or request timeout unless configured | connectTimeout on the client and timeout on each HttpRequest |
Node.js fetch | No overall request timeout | AbortSignal.timeout(ms), or a signal tied to the request deadline |
| gRPC (all languages) | No deadline | A deadline on every call, derived from the incoming one |
| Envoy route | 15 s route timeout; no per-try timeout | timeout and per_try_timeout that match the service budget |
| nginx proxy | proxy_connect_timeout and proxy_read_timeout 60 s | Values that sit just above the upstream service’s own budget |
| PostgreSQL | statement_timeout 0 (disabled) | A per-role or per-transaction limit below the service budget |
See the requests timeouts documentation, the Go context package, and Envoy’s timeout FAQ for details.
Proxies and meshes add their own layer. If a service mesh sits between services, its route timeout is another waiter in the chain and must follow the same nesting. A mesh timeout that is longer than the application’s client timeout does nothing; one that is shorter overrides the application without anyone noticing in code review.
How do you pick the numbers?
Pick them from two directions and make them meet. From the top, the edge budget comes from the product: how long a user or API client will reasonably wait. From the bottom, each per-try timeout comes from the dependency’s measured latency when it is healthy.
- Measure the dependency’s healthy tail. Use p99 or p99.9 over a normal week, per endpoint, not the average.
- Set the per-try timeout a little above that tail. Requests slower than this are more likely stuck than slow, and waiting longer rarely helps.
- Multiply by attempts and add backoff. That is the call’s cost to the caller’s budget.
- Check it fits. If the sum of calls exceeds the caller’s budget, the design needs to change, not the numbers.
| Hop | Healthy p99 | Per-try timeout | Attempts | Cost to caller |
|---|---|---|---|---|
| Gateway → Orders | 1.2 s | 8 s (the Orders budget) | 1 | 8 s of the 10 s edge budget |
| Orders → Inventory | 600 ms | 2 s | 2 | 4.3 s with backoff |
| Orders → Pricing | 300 ms | 1.5 s | 1 | 1.5 s |
| Inventory → PostgreSQL | 40 ms | 1 s statement timeout | 1 | 1 s of the 2 s Inventory attempt |
Illustrative numbers. The pattern matters more than the values: every row fits inside the row above it.
If a dependency’s healthy p99 is already close to the caller’s whole budget, no timeout value fixes it. The options are architectural: cache the result, precompute it, call it in parallel with other work, return a partial response without it, or move it off the synchronous path to a queue. A timeout budget that doesn’t add up is a design review finding, not a tuning task.
How do you propagate a deadline across services?
Propagate the remaining time, not a fixed timeout. A deadline is the moment by which the whole request must finish; a timeout is a duration for one hop. When each service derives its outbound timeouts from the remaining time on the incoming request, nesting happens automatically: a request that spent 7 seconds before reaching you only gets 3 seconds from you, whatever your static configuration says.
gRPC: deadlines are built in
gRPC carries the client’s deadline to the server in the grpc-timeout header, and the server sees it on the call’s context. If the server uses that incoming context for its own outbound calls, the deadline propagates with no extra code. In Go that means passing the handler’s ctx through; grpc-java propagates the deadline through io.grpc.Context. The gRPC deadlines guide recommends setting a deadline on every call, because the default is none.
HTTP: pick a header and use it everywhere
HTTP has no standard deadline header, so services agree on one. Send the remaining duration (for example, milliseconds left), not an absolute timestamp, so clock skew between hosts does not matter. Some proxies help: Envoy, for instance, can pass its expected timeout upstream in the x-envoy-expected-rq-timeout-ms header. Whatever the convention, it has to be read on the way in and written on the way out, in every service.
The most common way propagation breaks is small: a handler starts a fresh context or uses a default client, and the chain loses its deadline at that hop.
func (h *Handler) Reserve(w http.ResponseWriter, r *http.Request) {
ctx := context.Background() // drops the caller's deadline
req, _ := http.NewRequestWithContext(ctx, http.MethodPost, inventoryURL, r.Body)
resp, err := http.DefaultClient.Do(req) // DefaultClient: no timeout
if err != nil {
http.Error(w, err.Error(), http.StatusBadGateway)
return
}
defer resp.Body.Close()
io.Copy(w, resp.Body)
}
The fix has two halves. On the way in, turn the caller’s remaining budget into a context deadline:
const deadlineHeader = "X-Deadline-Ms" // a convention your services agree on
func WithDeadline(defaultBudget time.Duration, next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
budget := defaultBudget
if v := r.Header.Get(deadlineHeader); v != "" {
if ms, err := strconv.ParseInt(v, 10, 64); err == nil {
if ms <= 0 {
http.Error(w, "deadline exceeded", http.StatusGatewayTimeout)
return // caller already gave up: do no work
}
if d := time.Duration(ms) * time.Millisecond; d < budget {
budget = d
}
}
}
ctx, cancel := context.WithTimeout(r.Context(), budget)
defer cancel()
next.ServeHTTP(w, r.WithContext(ctx))
})
}
On the way out, cap each call by both its per-try limit and the remaining deadline, and pass a slightly smaller budget downstream so the callee gives up before you do:
var dialer = &net.Dialer{Timeout: 300 * time.Millisecond}
var client = &http.Client{Transport: &http.Transport{
DialContext: dialer.DialContext,
TLSHandshakeTimeout: 300 * time.Millisecond,
ResponseHeaderTimeout: 3 * time.Second,
}}
func ReserveStock(ctx context.Context, body []byte) (Reservation, error) {
var out Reservation
// Per-try cap. Never later than the parent deadline.
ctx, cancel := context.WithTimeout(ctx, 3*time.Second)
defer cancel()
dl, _ := ctx.Deadline()
// Keep 50 ms to send our own response; fail fast if too little is left.
remaining := time.Until(dl) - 50*time.Millisecond
if remaining < 100*time.Millisecond {
return out, context.DeadlineExceeded
}
req, err := http.NewRequestWithContext(ctx, http.MethodPost, inventoryURL,
bytes.NewReader(body))
if err != nil {
return out, err
}
req.Header.Set(deadlineHeader, strconv.FormatInt(remaining.Milliseconds(), 10))
resp, err := client.Do(req)
if err != nil {
return out, err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return out, fmt.Errorf("inventory: %s", resp.Status)
}
// Read the body here, before the deferred cancel() runs.
return out, json.NewDecoder(resp.Body).Decode(&out)
}
context.WithTimeout never extends a parent’s deadline: if the parent expires sooner, the child expires with it. That property is what makes the per-try cap safe to write as a constant.
What should a service do when the deadline expires?
Stop working, stop calling, and say so clearly. A deadline that is propagated but ignored only moves the orphaned work around.
- Check before expensive work. Before a query, a downstream call, or a heavy computation, check the time left. If it is less than the work needs, fail now.
- Count queue time. A request that waited in a server queue or thread pool for most of its budget should be dropped when it is dequeued, not processed.
- Cancel downstream work. Pass the context to database and client calls so they stop when it expires. In Go,
database/sqlpasses the context to the driver, and drivers that support cancellation stop the query. - Set a database backstop. A role-level
statement_timeoutcatches any path that forgot the context. - Return a distinct, non-retryable error. A gRPC
DEADLINE_EXCEEDEDor an HTTP 504 tells the caller the budget is gone. Retrying it is pointless, since the caller’s own deadline has passed too.
-- Applies to new sessions for this role; individual transactions can lower it
ALTER ROLE inventory_app SET statement_timeout = '2s';
-- Inside a transaction, derive a tighter limit from the request's remaining time
SET LOCAL statement_timeout = '800ms';
A timeout protects the caller. A propagated deadline protects everyone below it.
How do retries and fan-out fit inside the budget?
Retries spend the budget in series; fan-out spends it in parallel. Both have to be counted.
Retries. Each attempt needs its own per-try timeout, and all attempts plus backoff must fit in the call’s share of the budget. Before each retry, check the remaining deadline; a retry that cannot finish in time is load with no possible benefit. At a proxy, keep per_try_timeout × attempts under the route timeout. Where to retry, how much, and with what backoff is covered in the retry storm post, and Circuit Breaker vs. Retry vs. Load Shedding compares retries with the other controls.
routes:
- match: { prefix: "/orders" }
route:
cluster: orders
timeout: 10s # the edge budget for this route
retry_policy:
retry_on: "connect-failure,refused-stream"
num_retries: 1
per_try_timeout: 4s # 2 attempts x 4 s fits inside 10 s
Fan-out. When a service calls several dependencies, the shape of the calls decides how the budget is split. Sequential calls divide it; parallel calls share it.
For sequential calls, give each call a share and derive its timeout from the time left when it starts, so a slow first call shrinks the later ones instead of pushing the total past the deadline. For parallel calls, all of them inherit the same deadline; if some results are optional, stop waiting for those at a shorter internal cutoff and return a partial response. Large fan-out also raises the odds that some call hits its tail, which is why the per-try timeout should be set from p99 or p99.9, not the median.
A step-by-step method for setting timeouts across a chain
Use this when designing a new call path, or when auditing an existing one after an incident.
- Set the edge budget. Decide how long the client or user will wait for this operation, and configure it at the outermost layer.
- Map the call graph. List every synchronous hop the request makes, including proxies, meshes, and database calls, and whether calls are sequential or parallel.
- Measure healthy latency per hop. Collect p99 or p99.9 for each dependency endpoint over a normal period.
- Assign per-try timeouts and attempts. Set each per-try timeout a little above the healthy tail, and decide how many attempts, if any, that hop may make.
- Check the budget adds up. Verify that attempts, backoff, other calls, own work, and a margin fit inside each caller’s budget, from the edge down.
- Propagate the deadline. Pass the remaining time on every hop with gRPC deadlines or an agreed header, and derive outbound timeouts from it.
- Enforce it in the callee. Check the deadline before expensive work, drop requests that expired in a queue, cancel downstream calls, and set a database statement timeout as a backstop.
- Test with injected latency. Slow down one dependency in a test environment and confirm that inner hops give up first and no work continues after the edge returns.
How do you review timeouts in a pull request?
Timeout bugs rarely look like bugs in a diff. A new client with default settings, a context replaced with a fresh one, or a mesh route changed in another repository all pass tests. The review questions below catch most of them; for the broader practice see How to Review a Pull Request for Production Reliability Risks.
- Has a connect and an overall timeout
- Uses the incoming context or deadline
- Is not on a default client with no limit
- Timeout × attempts + backoff fits the caller’s budget
- Proxy and mesh timeouts match the code
- Database statement timeout below the request budget
- Remaining time checked before expensive work
- Downstream calls cancelled
- Deadline errors are not retried
Many of these questions cannot be answered from the diff alone. The caller’s budget lives in another service; the mesh timeout lives in a deployment repository; the shared HTTP client lives in a platform package. That cross-file context is what makes timeout review hard, and what How to Assess the Blast Radius of a Code Change is about. Developers ask variations of this regularly; a Stack Overflow search for timeout chains and deadline propagation shows how often the answer is “it depends on the caller’s timeout”. The Google SRE book chapter on addressing cascading failures covers deadline propagation as one of the core defenses.
How Tomosu helps
Tomosu analyzes a repository and each pull request for production reliability risk, including the timeout and deadline patterns described here:
- Outbound calls without timeouts, including clients created with library defaults that never time out.
- Dropped deadlines: handlers that start a fresh context or ignore the incoming one before calling downstream.
- Retry and timeout combinations where per-try timeout × attempts exceeds the budget the caller allows.
- Blast radius: which request paths reach a changed client or dependency, so a missing timeout on a hot path is weighed differently from one in a batch job.
These findings feed the Production Reliability Index. To see them on your own services, start with the repository scan guide.
Scan your repository with Tomosu →
Key takeaways
- Timeouts across a chain must be nested: each hop gives up before its caller does.
- Inverted timeouts create orphaned work below the edge and amplify retries.
- Budget from the edge inward: attempts × per-try timeout + backoff + other calls + own work + margin must fit.
- Set per-try timeouts from a dependency’s healthy p99 or p99.9, not from guesses or averages.
- Propagate the remaining time, with gRPC deadlines or an agreed header, and derive every outbound timeout from it.
- When the deadline expires, stop work, cancel downstream calls, and return a non-retryable error.
- If the budget cannot add up, change the design: cache, parallelize, degrade, or go asynchronous.
Frequently asked questions
How should timeouts be set across a chain of microservices?
Set them from the edge inward so each hop gives up before its caller does. Start from the user-facing budget, subtract each hop’s own work and a margin, and give each downstream call a per-try timeout based on the dependency’s healthy tail latency. Propagate the remaining time as a deadline so every service derives its outbound timeouts from it.
Should an inner service have a shorter timeout than its caller?
Yes. If an inner service waits longer than its caller, the caller returns an error while the inner service keeps working on a request nobody will use. The inner timeout, multiplied by the number of attempts and plus any backoff, should fit inside the caller’s remaining budget.
What is deadline propagation?
Deadline propagation means passing the time a request has left from one service to the next, so each service knows how long it may spend. gRPC does this with the grpc-timeout header and the call context. Over HTTP, services agree on a header that carries the remaining milliseconds and derive their outbound timeouts from it.
What is the difference between a timeout and a deadline?
A timeout is a duration for one call, such as 2 seconds to wait for a dependency. A deadline is the point in time by which the whole request must finish. Deadlines compose across a chain because each hop can compute how much time is left; fixed timeouts do not, because each is set without knowing how much time was already spent.
How do I choose a timeout value for a service call?
Measure the dependency’s healthy p99 or p99.9 latency for that endpoint and set the per-try timeout a little above it. Then check that the attempts, backoff, and other work fit inside the caller’s budget. If the dependency’s normal tail already approaches the whole budget, change the design rather than the number.
Do HTTP clients have a timeout by default?
Often not. Python requests has no timeout unless you pass one, a Go http.Client with a zero Timeout never times out, Java’s java.net.http.HttpClient has no request timeout unless configured, Node.js fetch has no overall request timeout, and gRPC calls have no deadline by default. Every outbound call should set one explicitly.
How do retries affect timeouts in a microservice chain?
Each retry spends more of the caller’s budget. The per-try timeout multiplied by the number of attempts, plus backoff, must fit inside the time the caller allows, or the caller gives up while retries are still running. Check the remaining deadline before each retry and do not retry deadline-exceeded errors.
A timeout budget only works if every hop honors it, and the hop that doesn’t is usually one line in a file nobody reviewed. Tomosu finds those lines across the repository. Assess your repository →