Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · API timeouts

How to Set Timeouts Across a Chain of Microservices

Tomosu AI·15 min read·

Every service in the chain has a timeout, and each one was chosen by a different team at a different time. The gateway gives up after 10 seconds, the service behind it waits 30, and the one behind that waits forever. When a dependency slows down, those numbers decide whether the slowdown stays local or spreads.

Quick answer

Set timeouts across a chain of microservices from the edge inward: each hop must give up before its caller does. Start from the user-facing budget, subtract each hop’s own work and a margin, and pass the remaining time downstream as a deadline. Every callee checks that deadline before it starts work and stops when it expires.

The previous post in this cluster, API Timeouts After Deployment: A Practical Triage Checklist, is about restoring service once timeouts are firing. This one is about the design that keeps a slow dependency from turning into a slow system: how to choose the numbers, how to make them agree across services, and how to check them in review.

Why do timeouts across a microservice chain go wrong?

They go wrong because each timeout is set locally, but the failure is global. A timeout limits how long one caller waits for one callee. In a chain, the timeouts only work if they are nested: each hop has to finish, or give up, before the hop above it stops waiting. When they are not nested, the chain has inverted timeouts, and the inner services keep working long after the outer ones have returned an error.

THE SAME CHAIN, TWO WAYS TO SET TIMEOUTS INVERTED: INNER TIMEOUTS LONGER THAN OUTER GATEWAYORDERSINVENTORYDATABASE timeout 10 stimeout 30 stimeout 60 sno limit Gateway returns 504 at 10 s; Orders and Inventory keep working up to 60 s for a caller that left. BUDGETED: EACH HOP SHORTER THAN ITS CALLER GATEWAYORDERSINVENTORYDATABASE timeout 10 stimeout 8 stimeout 5 sstatement 2 s Each hop gives up before its caller, and cancellation stops the work below it.
Inverted timeouts turn every slow request into orphaned work below the edge. Nested timeouts stop that work when the user stops waiting.

Inverted timeouts cause three problems at once:

The opposite mistake is just as common: timeouts so short that a healthy dependency’s normal tail latency trips them, producing errors and retries on a good day. The goal is not “short” or “long”. It is nested and measured.

What is a timeout budget?

A timeout budget is the total time a request is allowed to take, divided among the hops and attempts that serve it. It starts at the edge, where the product decides how long a user or client will wait, and shrinks at every hop by the time that hop has already spent.

For any hop, the rule is simple arithmetic: the time you allow a downstream call, multiplied by the attempts you allow, plus backoff, plus any other calls and your own work, plus a margin to send the response back, must be no more than the time your caller gave you.

ORDERS HAS 8 S FROM THE GATEWAY. WHERE DOES IT GO? Inventory try 1 · 2 s Inventory try 2 · 2 s Pricing · 1.5 s Margin · 1.5 s own work 0.7 s backoff 0.3 s 0 s2 s4 s6 s8 s attempts × per-try timeout + backoff + other calls + own work + margin ≤ remaining budget 2 × 2 s + 0.3 s + 1.5 s + 0.7 s + 1.5 s = 8.0 s
A budget turns “how long should this timeout be?” into arithmetic that reviewers can check.

Two consequences fall out of the arithmetic. First, retries are not free: two attempts at 2 seconds cost 4 seconds of the caller’s budget, so adding a retry without shortening the per-try timeout silently breaks the nesting. Second, the budget runs out faster the deeper the chain. A five-hop chain with a 2-second edge budget leaves very little for the last hop, which is a design signal in itself: the deepest calls need to be the fastest, cached, or moved off the request path.

Which timeouts does each call actually need?

“The timeout” is usually several settings. A call can hang while connecting, during a TLS handshake, while waiting for response headers, or while reading a slow body. A read timeout alone does not bound the total: a server that sends one byte every few seconds can keep a call alive indefinitely. Each outbound call needs at least:

The defaults are the trap. Many common clients do not time out unless you tell them to:

Client or layerDefaultWhat to set
Python requestsNo timeout unless you pass timeouttimeout=(connect, read) on every call, or a session wrapper that enforces it
Go net/httphttp.Client with zero Timeout never times out; http.DefaultClient is oneA context deadline per request, plus dial, TLS, and response header timeouts on the transport
Java java.net.http.HttpClientNo connect or request timeout unless configuredconnectTimeout on the client and timeout on each HttpRequest
Node.js fetchNo overall request timeoutAbortSignal.timeout(ms), or a signal tied to the request deadline
gRPC (all languages)No deadlineA deadline on every call, derived from the incoming one
Envoy route15 s route timeout; no per-try timeouttimeout and per_try_timeout that match the service budget
nginx proxyproxy_connect_timeout and proxy_read_timeout 60 sValues that sit just above the upstream service’s own budget
PostgreSQLstatement_timeout 0 (disabled)A per-role or per-transaction limit below the service budget

See the requests timeouts documentation, the Go context package, and Envoy’s timeout FAQ for details.

Proxies and meshes add their own layer. If a service mesh sits between services, its route timeout is another waiter in the chain and must follow the same nesting. A mesh timeout that is longer than the application’s client timeout does nothing; one that is shorter overrides the application without anyone noticing in code review.

How do you pick the numbers?

Pick them from two directions and make them meet. From the top, the edge budget comes from the product: how long a user or API client will reasonably wait. From the bottom, each per-try timeout comes from the dependency’s measured latency when it is healthy.

  1. Measure the dependency’s healthy tail. Use p99 or p99.9 over a normal week, per endpoint, not the average.
  2. Set the per-try timeout a little above that tail. Requests slower than this are more likely stuck than slow, and waiting longer rarely helps.
  3. Multiply by attempts and add backoff. That is the call’s cost to the caller’s budget.
  4. Check it fits. If the sum of calls exceeds the caller’s budget, the design needs to change, not the numbers.
HopHealthy p99Per-try timeoutAttemptsCost to caller
Gateway → Orders1.2 s8 s (the Orders budget)18 s of the 10 s edge budget
Orders → Inventory600 ms2 s24.3 s with backoff
Orders → Pricing300 ms1.5 s11.5 s
Inventory → PostgreSQL40 ms1 s statement timeout11 s of the 2 s Inventory attempt

Illustrative numbers. The pattern matters more than the values: every row fits inside the row above it.

When the numbers don’t fit

If a dependency’s healthy p99 is already close to the caller’s whole budget, no timeout value fixes it. The options are architectural: cache the result, precompute it, call it in parallel with other work, return a partial response without it, or move it off the synchronous path to a queue. A timeout budget that doesn’t add up is a design review finding, not a tuning task.

How do you propagate a deadline across services?

Propagate the remaining time, not a fixed timeout. A deadline is the moment by which the whole request must finish; a timeout is a duration for one hop. When each service derives its outbound timeouts from the remaining time on the incoming request, nesting happens automatically: a request that spent 7 seconds before reaching you only gets 3 seconds from you, whatever your static configuration says.

PASSING THE REMAINING BUDGET DOWN THE CHAIN Gateway Orders Inventory Database budget 10 s X-Deadline-Ms: 9800 X-Deadline-Ms: 3000 statement_timeout 2s own work 0.6 s per-try cap (3 s) < 9.1 s left rows 200 OK 200 OK in 4.1 s If the time left is below what a call needs, fail fast with DEADLINE_EXCEEDED or 504. Don’t call.
Each hop sends less than it received. The callee’s budget is always shorter than the caller’s wait, so the callee gives up first.

gRPC: deadlines are built in

gRPC carries the client’s deadline to the server in the grpc-timeout header, and the server sees it on the call’s context. If the server uses that incoming context for its own outbound calls, the deadline propagates with no extra code. In Go that means passing the handler’s ctx through; grpc-java propagates the deadline through io.grpc.Context. The gRPC deadlines guide recommends setting a deadline on every call, because the default is none.

HTTP: pick a header and use it everywhere

HTTP has no standard deadline header, so services agree on one. Send the remaining duration (for example, milliseconds left), not an absolute timestamp, so clock skew between hosts does not matter. Some proxies help: Envoy, for instance, can pass its expected timeout upstream in the x-envoy-expected-rq-timeout-ms header. Whatever the convention, it has to be read on the way in and written on the way out, in every service.

The most common way propagation breaks is small: a handler starts a fresh context or uses a default client, and the chain loses its deadline at that hop.

orders/reserve.godeadline dropped, no timeout
func (h *Handler) Reserve(w http.ResponseWriter, r *http.Request) {
	ctx := context.Background()               // drops the caller's deadline
	req, _ := http.NewRequestWithContext(ctx, http.MethodPost, inventoryURL, r.Body)
	resp, err := http.DefaultClient.Do(req)   // DefaultClient: no timeout
	if err != nil {
		http.Error(w, err.Error(), http.StatusBadGateway)
		return
	}
	defer resp.Body.Close()
	io.Copy(w, resp.Body)
}

The fix has two halves. On the way in, turn the caller’s remaining budget into a context deadline:

platform/deadline.goserver side
const deadlineHeader = "X-Deadline-Ms"   // a convention your services agree on

func WithDeadline(defaultBudget time.Duration, next http.Handler) http.Handler {
	return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
		budget := defaultBudget
		if v := r.Header.Get(deadlineHeader); v != "" {
			if ms, err := strconv.ParseInt(v, 10, 64); err == nil {
				if ms <= 0 {
					http.Error(w, "deadline exceeded", http.StatusGatewayTimeout)
					return   // caller already gave up: do no work
				}
				if d := time.Duration(ms) * time.Millisecond; d < budget {
					budget = d
				}
			}
		}
		ctx, cancel := context.WithTimeout(r.Context(), budget)
		defer cancel()
		next.ServeHTTP(w, r.WithContext(ctx))
	})
}

On the way out, cap each call by both its per-try limit and the remaining deadline, and pass a slightly smaller budget downstream so the callee gives up before you do:

orders/inventory_client.goclient side
var dialer = &net.Dialer{Timeout: 300 * time.Millisecond}
var client = &http.Client{Transport: &http.Transport{
	DialContext:           dialer.DialContext,
	TLSHandshakeTimeout:   300 * time.Millisecond,
	ResponseHeaderTimeout: 3 * time.Second,
}}

func ReserveStock(ctx context.Context, body []byte) (Reservation, error) {
	var out Reservation
	// Per-try cap. Never later than the parent deadline.
	ctx, cancel := context.WithTimeout(ctx, 3*time.Second)
	defer cancel()
	dl, _ := ctx.Deadline()
	// Keep 50 ms to send our own response; fail fast if too little is left.
	remaining := time.Until(dl) - 50*time.Millisecond
	if remaining < 100*time.Millisecond {
		return out, context.DeadlineExceeded
	}
	req, err := http.NewRequestWithContext(ctx, http.MethodPost, inventoryURL,
		bytes.NewReader(body))
	if err != nil {
		return out, err
	}
	req.Header.Set(deadlineHeader, strconv.FormatInt(remaining.Milliseconds(), 10))
	resp, err := client.Do(req)
	if err != nil {
		return out, err
	}
	defer resp.Body.Close()
	if resp.StatusCode != http.StatusOK {
		return out, fmt.Errorf("inventory: %s", resp.Status)
	}
	// Read the body here, before the deferred cancel() runs.
	return out, json.NewDecoder(resp.Body).Decode(&out)
}

context.WithTimeout never extends a parent’s deadline: if the parent expires sooner, the child expires with it. That property is what makes the per-try cap safe to write as a constant.

What should a service do when the deadline expires?

Stop working, stop calling, and say so clearly. A deadline that is propagated but ignored only moves the orphaned work around.

PostgreSQL · backstop per rolebounded queries
-- Applies to new sessions for this role; individual transactions can lower it
ALTER ROLE inventory_app SET statement_timeout = '2s';

-- Inside a transaction, derive a tighter limit from the request's remaining time
SET LOCAL statement_timeout = '800ms';

A timeout protects the caller. A propagated deadline protects everyone below it.

How do retries and fan-out fit inside the budget?

Retries spend the budget in series; fan-out spends it in parallel. Both have to be counted.

Retries. Each attempt needs its own per-try timeout, and all attempts plus backoff must fit in the call’s share of the budget. Before each retry, check the remaining deadline; a retry that cannot finish in time is load with no possible benefit. At a proxy, keep per_try_timeout × attempts under the route timeout. Where to retry, how much, and with what backoff is covered in the retry storm post, and Circuit Breaker vs. Retry vs. Load Shedding compares retries with the other controls.

Envoy route · attempts fit inside the route timeoutnested
routes:
- match: { prefix: "/orders" }
  route:
    cluster: orders
    timeout: 10s                       # the edge budget for this route
    retry_policy:
      retry_on: "connect-failure,refused-stream"
      num_retries: 1
      per_try_timeout: 4s              # 2 attempts x 4 s fits inside 10 s

Fan-out. When a service calls several dependencies, the shape of the calls decides how the budget is split. Sequential calls divide it; parallel calls share it.

SEQUENTIAL CALLS SPLIT THE BUDGET; PARALLEL CALLS SHARE IT deadline 6 s SEQUENTIAL One after another Profile 1.5 s Pricing 2 s Stock 1.5 s Total ≈ sum = 5 s. Each call only gets what the earlier ones left. PARALLEL FAN-OUT ProfilePricingStock 1.5 s 2 s 1.5 s 0 s2 s4 s6 s Parallel: total ≈ slowest call = 2 s, and every call can use the full remaining deadline.
In a sequential chain, a slow first call steals budget from every call after it. In fan-out, the slowest call sets the total.

For sequential calls, give each call a share and derive its timeout from the time left when it starts, so a slow first call shrinks the later ones instead of pushing the total past the deadline. For parallel calls, all of them inherit the same deadline; if some results are optional, stop waiting for those at a shorter internal cutoff and return a partial response. Large fan-out also raises the odds that some call hits its tail, which is why the per-try timeout should be set from p99 or p99.9, not the median.

A step-by-step method for setting timeouts across a chain

Use this when designing a new call path, or when auditing an existing one after an incident.

  1. Set the edge budget. Decide how long the client or user will wait for this operation, and configure it at the outermost layer.
  2. Map the call graph. List every synchronous hop the request makes, including proxies, meshes, and database calls, and whether calls are sequential or parallel.
  3. Measure healthy latency per hop. Collect p99 or p99.9 for each dependency endpoint over a normal period.
  4. Assign per-try timeouts and attempts. Set each per-try timeout a little above the healthy tail, and decide how many attempts, if any, that hop may make.
  5. Check the budget adds up. Verify that attempts, backoff, other calls, own work, and a margin fit inside each caller’s budget, from the edge down.
  6. Propagate the deadline. Pass the remaining time on every hop with gRPC deadlines or an agreed header, and derive outbound timeouts from it.
  7. Enforce it in the callee. Check the deadline before expensive work, drop requests that expired in a queue, cancel downstream calls, and set a database statement timeout as a backstop.
  8. Test with injected latency. Slow down one dependency in a test environment and confirm that inner hops give up first and no work continues after the edge returns.

How do you review timeouts in a pull request?

Timeout bugs rarely look like bugs in a diff. A new client with default settings, a context replaced with a fresh one, or a mesh route changed in another repository all pass tests. The review questions below catch most of them; for the broader practice see How to Review a Pull Request for Production Reliability Risks.

Every outbound call
  • Has a connect and an overall timeout
  • Uses the incoming context or deadline
  • Is not on a default client with no limit
Nesting
  • Timeout × attempts + backoff fits the caller’s budget
  • Proxy and mesh timeouts match the code
  • Database statement timeout below the request budget
On expiry
  • Remaining time checked before expensive work
  • Downstream calls cancelled
  • Deadline errors are not retried

Many of these questions cannot be answered from the diff alone. The caller’s budget lives in another service; the mesh timeout lives in a deployment repository; the shared HTTP client lives in a platform package. That cross-file context is what makes timeout review hard, and what How to Assess the Blast Radius of a Code Change is about. Developers ask variations of this regularly; a Stack Overflow search for timeout chains and deadline propagation shows how often the answer is “it depends on the caller’s timeout”. The Google SRE book chapter on addressing cascading failures covers deadline propagation as one of the core defenses.

How Tomosu helps

Tomosu analyzes a repository and each pull request for production reliability risk, including the timeout and deadline patterns described here:

These findings feed the Production Reliability Index. To see them on your own services, start with the repository scan guide.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

How should timeouts be set across a chain of microservices?

Set them from the edge inward so each hop gives up before its caller does. Start from the user-facing budget, subtract each hop’s own work and a margin, and give each downstream call a per-try timeout based on the dependency’s healthy tail latency. Propagate the remaining time as a deadline so every service derives its outbound timeouts from it.

Should an inner service have a shorter timeout than its caller?

Yes. If an inner service waits longer than its caller, the caller returns an error while the inner service keeps working on a request nobody will use. The inner timeout, multiplied by the number of attempts and plus any backoff, should fit inside the caller’s remaining budget.

What is deadline propagation?

Deadline propagation means passing the time a request has left from one service to the next, so each service knows how long it may spend. gRPC does this with the grpc-timeout header and the call context. Over HTTP, services agree on a header that carries the remaining milliseconds and derive their outbound timeouts from it.

What is the difference between a timeout and a deadline?

A timeout is a duration for one call, such as 2 seconds to wait for a dependency. A deadline is the point in time by which the whole request must finish. Deadlines compose across a chain because each hop can compute how much time is left; fixed timeouts do not, because each is set without knowing how much time was already spent.

How do I choose a timeout value for a service call?

Measure the dependency’s healthy p99 or p99.9 latency for that endpoint and set the per-try timeout a little above it. Then check that the attempts, backoff, and other work fit inside the caller’s budget. If the dependency’s normal tail already approaches the whole budget, change the design rather than the number.

Do HTTP clients have a timeout by default?

Often not. Python requests has no timeout unless you pass one, a Go http.Client with a zero Timeout never times out, Java’s java.net.http.HttpClient has no request timeout unless configured, Node.js fetch has no overall request timeout, and gRPC calls have no deadline by default. Every outbound call should set one explicitly.

How do retries affect timeouts in a microservice chain?

Each retry spends more of the caller’s budget. The per-try timeout multiplied by the number of attempts, plus backoff, must fit inside the time the caller allows, or the caller gives up while retries are still running. Check the remaining deadline before each retry and do not retry deadline-exceeded errors.


A timeout budget only works if every hop honors it, and the hop that doesn’t is usually one line in a file nobody reviewed. Tomosu finds those lines across the repository. Assess your repository →