The release went out, and a few minutes later the dashboards turned red with 504s and client timeouts. You need two answers quickly: whether to undo the deploy, and where the extra time is going. This checklist gives you both, in that order.
When API timeouts start right after a deployment, mitigate first and diagnose second. Confirm the timeouts line up with the rollout, roll back or turn off the feature flag if that is safe, then find where the time is spent. The component that reports the timeout is rarely the one that is slow. Most post-deploy timeouts fall into five groups:
- New code path: an added query, N+1 loop, or remote call on a hot endpoint.
- Resource starvation: CPU throttling, GC pressure, or fewer ready replicas.
- Contention: a pool, database lock, or blocked event loop that requests queue behind.
- Shared cause: a migration, dependency, or traffic change that hits old and new pods alike.
- Config drift: changed timeouts, endpoints, environment variables, or probes.
Raising the timeout is the tempting fix, and it is almost never the right first move. A longer timeout keeps more requests in flight, holds more threads and connections, and turns a partial outage into a full one. This guide walks through the triage in the order that restores service fastest, then shows how to tell the five causes apart with evidence rather than guesses.
What does an API timeout after deployment actually tell you?
An API timeout means some component stopped waiting for a response before one arrived. That is all it means. A timeout is a limit on how long one hop will wait; it is reported by the waiter, not by the component that was slow. In a typical request path there are three or four waiters, each with its own limit, and whichever limit is shortest decides who reports the error.
The same incident can therefore surface as very different messages depending on where you look:
# nginx error log (the client gets 504 Gateway Timeout)
upstream timed out (110: Connection timed out) while reading response header from upstream
# Envoy / Istio: 504 with this response body, response flag UT in access logs
upstream request timeout
# Go service calling a dependency
Get "http://inventory:8080/stock": context deadline exceeded
# Python client
requests.exceptions.ReadTimeout: HTTPConnectionPool(host='payments', port=80): Read timed out. (read timeout=0.5)
Each of these tells you who gave up. None of them tells you where the time went. That distinction drives the whole triage: you start at the layer that reported the timeout and walk inward until you find the hop whose own work, or whose own wait, grew after the release.
Notice the wasted work at the end. When a caller gives up and the callee does not, the callee keeps spending capacity on answers nobody will read. That is how a post-deploy slowdown feeds itself, and why the second post in this cluster, How to Set Timeouts Across a Chain of Microservices, focuses on deadlines that shrink as they travel inward.
Should you roll back before you diagnose?
Usually yes. If the timeouts started with the rollout and rolling back is safe, rolling back is the fastest way to restore service, and you can diagnose the bad version later from the evidence you captured. The goal of the first ten minutes is not to understand the bug. It is to stop user impact without destroying the evidence you will need.
Three conditions make rollback the right call: the timeouts began within minutes of the rollout, the new version is the only meaningful change in that window, and the previous version can run against the current data and schema. If a migration already dropped a column the old code reads, or the new version wrote data the old version cannot parse, rollback becomes its own incident. In that case, turning off a feature flag, scaling out, or shedding the new endpoint may be safer. The full decision is covered in How to Decide Whether to Roll Back a Release.
Before you roll back, spend sixty seconds capturing what will disappear with the bad pods:
# Which revision is live, and when did it start rolling out?
kubectl rollout history deployment/orders-api
kubectl get pods -l app=orders-api -o wide --sort-by=.metadata.creationTimestamp
# Snapshot resource use and recent logs from one new pod before it goes away
kubectl top pods -l app=orders-api
kubectl logs orders-api-7c9f8d6b5-x2x4q --since=15m > new-pod.log
# Thread / goroutine / heap dump if your runtime supports it (JVM example)
kubectl exec orders-api-7c9f8d6b5-x2x4q -- jcmd 1 Thread.print > threads.txt
# Roll back to the previous revision and watch it complete
kubectl rollout undo deployment/orders-api
kubectl rollout status deployment/orders-api
A longer timeout does not make the slow hop faster. It lets more requests pile up behind it, each holding a thread, a connection, and memory, so the service reaches its limits sooner. Raise a timeout only when evidence shows the work is legitimately slower and the capacity to carry it exists.
Is the deployment really the cause?
A deploy is the most visible change in any incident window, so it gets blamed first. Often it deserves to be, but check before you act on that assumption. The fastest test is to compare the new version with the old one during the rollout itself, while both are serving traffic.
- Split latency and errors by version. Most metrics pipelines can group by a pod label or image tag. If only new-version pods time out, the release is implicated. If old pods time out too, something shared changed.
- Line up the timestamps. The first timeout should follow the first new pod becoming ready, not precede it. A spike that started ten minutes before the rollout is not the rollout.
- List every other change in the window. Config pushes, feature flag flips, a dependency’s deploy, a database migration, a batch job, a traffic spike, or an infrastructure event.
- Check which requests fail. One endpoint, one tenant, or one region narrows the search more than any aggregate graph.
If you need to go further, Which Commit Caused the Production Incident? covers narrowing a release to the change that caused it, and How to Correlate Logs, Traces, and a Code Change During an Incident covers stitching the signals together. Developers hit the same ambiguity often; searches for 504 gateway timeout after deployment return many threads where the deploy was blamed for a load balancer or dependency change.
The API timeout triage checklist
Run these steps in order. Each one either narrows the cause or rules out a group of causes, and the order front-loads the checks that are cheapest and most decisive.
- Confirm and scope. Check that timeouts began with the rollout, and find which endpoints, pods, versions, and regions are affected.
- Mitigate. Roll back, disable the feature flag, or scale out, after capturing logs, dumps, and pod metrics from a new pod.
- Find the reporting layer and its limit. Identify which component gave up (client, gateway, service, or driver) and what its timeout value is.
- Read the latency distribution. Decide whether all requests got slower, a subset is pinned at a timeout value, or one group is slow.
- Check resources on new pods. Look at CPU throttling, memory and GC, restarts, and the number of ready replicas.
- Check for waiting. Look for requests queued on a connection pool, a database lock, a thread pool, or a blocked event loop.
- Trace one slow request. Compare a slow trace from the new version with a normal trace from the old one and find the span that grew or appeared.
- Diff the configuration, not just the code. Compare rendered environment variables, timeouts, endpoints, pool sizes, and probes between the two versions.
- New pods only, or all pods?
- One endpoint, or every endpoint?
- One region, zone, or tenant?
- Started with the rollout, not before?
- CPU throttling on new pods
- Memory, GC pauses, restarts
- Ready replicas vs desired
- Startup and warm-up time
- Connection pool wait time
- Database lock waits
- Thread pool or event loop lag
- Queue depth in front of workers
- Timeout values at each hop
- Dependency URLs and regions
- Pool sizes and limits
- Readiness probe and rollout strategy
Which of the five causes is it?
Four questions, asked in order, sort almost every post-deploy timeout into one of five groups. Each question is answerable from standard metrics and one or two traces.
1. Shared cause: old pods are slow too
If both versions time out, the release is probably not the cause, or it is only the trigger. Look for a schema migration that ran with the deploy, a dependency that deployed at the same time, a batch job, a traffic spike, or an infrastructure event. Migrations are the most common case that looks like a code problem: they ship with the release but act on the shared database.
2. Resource starvation on new pods
The new version needs more CPU or memory per request than the old one, or it has less to work with. Typical triggers are a heavier serialization library, a new per-request computation, larger in-memory caches, a changed CPU limit in the manifest, or a rollout that left fewer ready replicas. CPU throttling is easy to miss because CPU usage looks capped rather than high; the cAdvisor counters container_cpu_cfs_throttled_periods_total over container_cpu_cfs_periods_total show the throttled fraction.
3. Contention: requests wait behind something
Each request does roughly the same work as before, but spends much longer waiting for a shared resource: a database connection, a row or table lock, a worker thread, or the event loop in Node.js. Latency stays normal for requests that get through and jumps for those that queue. If connection pools are involved, HikariCP: Connection Is Not Available, Request Timed Out and Connection Leak vs. Pool Exhaustion go deeper on telling a leak from a long hold.
4. A new or slower code path
The release added work on a hot path: a query without an index, a loop that issues one query per item, a synchronous call to another service, or a larger response payload. A trace from the new version shows a span that was not there before, or a familiar span that is now much longer. The query side of this is covered in Why an API Is Slow in Production but Fast Locally.
5. Config drift
The code is fine but the environment around it changed: a timeout value in a new config file, a dependency URL pointing at another region, a pool size that fell back to its default when a key was renamed, or a readiness probe that now passes before the app is warm. Diff the rendered configuration of the two versions (environment, mounted config, Helm values after templating), not just the source code.
How do you read latency and timeout signals?
The shape of the latency distribution after the deploy is the single most informative signal you have. The average hides it; look at a histogram or at p50, p95, and p99 side by side.
| Signal after the deploy | What it usually means | Check next |
|---|---|---|
| p50 and p99 both rose | Every request does more work, or has less CPU to do it with | CPU throttling, GC time, a new span on the common path |
| p50 flat, p99 pinned at a round number (5 s, 15 s, 30 s) | A group of requests waits until some timer fires | Find which component has that timeout value; look for pool or lock waits behind it |
| Timeouts only on new-version pods | The release or its config | Trace diff between versions; rendered config diff |
| Timeouts on old and new pods alike | A shared change: migration, dependency, traffic | Database locks, dependency deploys, request rate |
| Worst in the first minutes, then easing | Warm-up: cold caches, JIT, connection setup, a probe passing too early | Readiness probe, cache hit rate, startup connection count |
| Getting steadily worse after the rollout finished | Something accumulates: a leak, a growing queue, a retry loop | Pool active count, queue depth, retry metrics, memory trend |
| 504 with a short duration | A hop with a shorter timeout than expected, or a connect failure | Gateway route timeouts, connect timeouts, DNS or endpoint changes |
| Only one endpoint or tenant affected | A code path or data shape specific to that traffic | Traces for that route; data volume for that tenant |
When failed requests all take almost exactly 5, 10, 15, 30, or 60 seconds, a timer somewhere is deciding their fate. Search your config for that value. The component that owns it is the one giving up, and whatever it is waiting on is the next hop to inspect. For example, nginx’s proxy_read_timeout defaults to 60 seconds, and an Envoy route timeout defaults to 15 seconds.
Which deploy-specific causes do people miss?
Some causes exist only because a deployment happened. They are easy to miss because the code diff looks harmless.
A migration waiting on a lock
In PostgreSQL, most ALTER TABLE forms need an ACCESS EXCLUSIVE lock. If a long transaction holds any lock on the table, the ALTER waits, and every query that arrives after it queues behind the ALTER, including plain reads. A quick migration becomes minutes of timeouts on every endpoint that touches the table. CREATE INDEX without CONCURRENTLY blocks writes to the table while it builds.
SELECT pid, pg_blocking_pids(pid) AS blocked_by,
state, wait_event_type,
now() - query_start AS waiting_for,
left(query, 70) AS query
FROM pg_stat_activity
WHERE cardinality(pg_blocking_pids(pid)) > 0
ORDER BY query_start;
ALTER TABLE orders ADD COLUMN fulfillment_region text;
CREATE INDEX orders_region_idx ON orders (fulfillment_region); -- blocks writes
SET lock_timeout = '3s'; -- fail fast instead of queueing traffic; retry later
ALTER TABLE orders ADD COLUMN fulfillment_region text;
-- run outside a transaction block
CREATE INDEX CONCURRENTLY orders_region_idx ON orders (fulfillment_region);
The PostgreSQL documentation describes lock_timeout and statement_timeout; both are worth setting for migration sessions.
A new call without a timeout
The most common post-deploy code cause is a new synchronous call to another service on a hot path. If the client has no timeout, one slow dependency holds a worker per request until something upstream gives up. Python’s requests, for example, never times out unless you pass timeout.
def get_order(order_id):
order = repo.load(order_id)
# added in this release: show estimated delivery on every order read
eta = requests.get(f"{SHIPPING_URL}/eta/{order_id}").json() # waits forever
return {**order.to_dict(), "eta": eta}
def get_order(order_id):
order = repo.load(order_id)
try:
resp = shipping.get(f"{SHIPPING_URL}/eta/{order_id}",
timeout=(0.2, 0.5)) # connect, read seconds
resp.raise_for_status()
eta = resp.json()
except requests.RequestException:
eta = None # ETA is optional; the order is not
return {**order.to_dict(), "eta": eta}
Capacity lost during the rollout
A rolling update removes old pods as new ones become ready. With the Kubernetes defaults of maxUnavailable: 25% and maxSurge: 25%, a service already running near its limit briefly has less capacity than it needs. If readiness probes pass before the application has warmed its caches, JIT, or connection pools, the new pods also start slower than the old ones they replace. See the Kubernetes Deployment documentation for the rolling update settings.
Cold caches and connection storms
A release that changes a cache key format, a serialization version, or a cache namespace starts with an empty cache, so every request goes to the database at once. That is the pattern in Cache Stampede: Why a Healthy Cache Can Overload Your Database. Similarly, many new pods opening full connection pools at the same moment can push a database to its connection limit.
The component that reports the timeout is the one that gave up. The fix belongs to the one that made it wait.
Mitigation and durable fix, by cause
| Cause | Mitigate now | Durable fix |
|---|---|---|
| New code path | Roll back or disable the flag for that path | Fix the query or loop, add a timeout and fallback to the new call, load test the route |
| Resource starvation | Scale out, or restore the previous resource limits | Profile the new per-request cost; set limits from measured usage |
| Contention (pool, lock, event loop) | Roll back; kill the blocking session if it is a migration | Shorter holds, no I/O inside transactions, lock_timeout on migrations, move CPU work off the event loop |
| Shared cause | Pause the job or migration; coordinate with the dependency owner | Separate migrations from releases; expand-and-contract schema changes |
| Config drift | Restore the previous config values | Review rendered config in the pull request; validate required keys at startup |
| Warm-up and rollout capacity | Pause the rollout; scale up before resuming | Readiness that reflects warm state; lower maxUnavailable; staged or canary rollout |
After the service is stable, ask what the release checks missed. A new call without a timeout, a migration without a lock timeout, or a resource limit change are all visible in the pull request before merge; they are rarely visible in unit tests. A PR Passed CI but Broke Production looks at that gap directly, and the Google SRE book chapter on managing incidents covers the coordination side of the first thirty minutes.
How Tomosu helps
Tomosu analyzes a repository and each pull request for production reliability risk. For post-deploy timeouts, the value is in seeing, before the release, the risks this checklist looks for after it:
- New outbound calls on request paths, and whether each has a timeout, a fallback, and a bounded retry.
- Blocking work in the wrong place: remote calls inside transactions, synchronous I/O on an event loop, and unbounded loops of queries.
- Changes outside the code that alter runtime behavior, such as timeout values, pool sizes, and resource limits in configuration and manifests.
- Blast radius: which endpoints and jobs share the code path or dependency that changed, so a risky change on a hot path is weighed accordingly.
These findings feed the Production Reliability Index, so a release that adds a slow dependency to a hot path is visible to reviewers before it ships. To try it on your own code, follow the repository scan guide.
Scan your repository with Tomosu →
Key takeaways
- A timeout tells you who gave up, not who was slow. Walk inward from the reporting layer.
- Mitigate first: if the new version is the only change and rollback is safe, roll back, after capturing evidence from a new pod.
- Split latency and errors by version. Old pods timing out too means a shared cause such as a migration or dependency.
- Read the distribution shape: a shift means more work, a spike at a round number means waiting, a second hump means one subset.
- Check deploy-only causes: migration locks, rollout capacity, early readiness, cold caches, and connection storms.
- Diff the rendered configuration, not just the code.
- Do not raise timeouts as a mitigation; it increases in-flight work and spreads the outage.
Frequently asked questions
Why do API timeouts start right after a deployment?
The most common reasons are a new or slower code path on a hot endpoint, less CPU or memory per request on the new pods, requests waiting on a connection pool or database lock, a shared change such as a migration that shipped with the release, or configuration drift such as changed timeouts or endpoints. Warm-up and reduced capacity during a rolling update can also cause short bursts of timeouts.
Should I roll back or debug first when timeouts follow a deploy?
Roll back first if the timeouts began with the rollout, the new version is the only meaningful change, and the previous version can run against the current schema and data. Capture logs, dumps, and pod metrics from a new pod before rolling back so you can diagnose the bad version afterwards.
Why does the gateway return 504 when the problem is in my service?
A 504 Gateway Timeout means the gateway or load balancer stopped waiting for your service. The gateway reports the error because its timeout was the shortest one on the path. The time was spent inside your service or one of its dependencies, so start from the gateway and trace inward.
Should I increase the timeout to stop the errors?
Not as a first response. A longer timeout lets more requests pile up behind the slow component, each holding a thread, connection, and memory, which usually makes the outage wider. Increase a timeout only when evidence shows the work is legitimately slower and there is capacity to carry the extra in-flight requests.
How can I tell if the deployment or something else caused the timeouts?
Split latency and errors by version during the rollout. If only new-version pods time out, the release or its configuration is implicated. If old pods time out too, look for a shared change in the same window, such as a database migration, a dependency deploy, a batch job, or a traffic spike.
Can a database migration cause API timeouts after a deploy?
Yes. In PostgreSQL, most ALTER TABLE forms need an ACCESS EXCLUSIVE lock. If the migration waits behind a long transaction, later queries on that table queue behind it, including reads. Setting lock_timeout on the migration session makes it fail fast instead of blocking traffic, and CREATE INDEX CONCURRENTLY avoids blocking writes while an index builds.
What does it mean when failed requests all take exactly 30 seconds?
A timer somewhere on the request path is set to 30 seconds and is deciding when those requests fail. Search your gateway, client, driver, and pool configuration for that value. The component that owns it is the one giving up, and the thing it is waiting on is the next hop to inspect.
Most post-deploy timeouts were visible in the pull request: a new call with no timeout, a migration with no lock timeout, a limit that changed. Tomosu surfaces those before the release. Assess your repository →