Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · API timeouts

API Timeouts After Deployment: A Practical Triage Checklist

Tomosu AI·14 min read·

The release went out, and a few minutes later the dashboards turned red with 504s and client timeouts. You need two answers quickly: whether to undo the deploy, and where the extra time is going. This checklist gives you both, in that order.

Quick answer

When API timeouts start right after a deployment, mitigate first and diagnose second. Confirm the timeouts line up with the rollout, roll back or turn off the feature flag if that is safe, then find where the time is spent. The component that reports the timeout is rarely the one that is slow. Most post-deploy timeouts fall into five groups:

Raising the timeout is the tempting fix, and it is almost never the right first move. A longer timeout keeps more requests in flight, holds more threads and connections, and turns a partial outage into a full one. This guide walks through the triage in the order that restores service fastest, then shows how to tell the five causes apart with evidence rather than guesses.

What does an API timeout after deployment actually tell you?

An API timeout means some component stopped waiting for a response before one arrived. That is all it means. A timeout is a limit on how long one hop will wait; it is reported by the waiter, not by the component that was slow. In a typical request path there are three or four waiters, each with its own limit, and whichever limit is shortest decides who reports the error.

The same incident can therefore surface as very different messages depending on where you look:

The same slowness, reported by different waiterssymptoms
# nginx error log (the client gets 504 Gateway Timeout)
upstream timed out (110: Connection timed out) while reading response header from upstream

# Envoy / Istio: 504 with this response body, response flag UT in access logs
upstream request timeout

# Go service calling a dependency
Get "http://inventory:8080/stock": context deadline exceeded

# Python client
requests.exceptions.ReadTimeout: HTTPConnectionPool(host='payments', port=80): Read timed out. (read timeout=0.5)

Each of these tells you who gave up. None of them tells you where the time went. That distinction drives the whole triage: you start at the layer that reported the timeout and walk inward until you find the hop whose own work, or whose own wait, grew after the release.

WHO REPORTS THE TIMEOUT VS WHERE THE TIME GOES CLIENT timeout 30 s GATEWAY route timeout 15 s reports the 504 API SERVICE (NEW) no deadline of its own DATABASE pool wait, slow SQL where time is spent Gateway Service waiting for the service 504 at 15 s waiting for a DB connection ~11 s query ~6 s finishes at 17 s, after the 504: wasted work 0 s5 s10 s15 s20 s Start where the timeout was reported, then walk inward to the hop whose work or wait grew after the release.
The gateway reports the 504, but the time was spent waiting for a database connection inside the new version of the service.

Notice the wasted work at the end. When a caller gives up and the callee does not, the callee keeps spending capacity on answers nobody will read. That is how a post-deploy slowdown feeds itself, and why the second post in this cluster, How to Set Timeouts Across a Chain of Microservices, focuses on deadlines that shrink as they travel inward.

Should you roll back before you diagnose?

Usually yes. If the timeouts started with the rollout and rolling back is safe, rolling back is the fastest way to restore service, and you can diagnose the bad version later from the evidence you captured. The goal of the first ten minutes is not to understand the bug. It is to stop user impact without destroying the evidence you will need.

THE FIRST 30 MINUTES AFTER TIMEOUTS START 0–5 MIN · CONFIRM Did timeouts beginwith the rollout?Which endpoints, pods,versions, regions? 5–10 MIN · MITIGATE Roll back if safe,or turn the flag offScale out if CPU-boundDon’t tune timeouts yet 10–30 MIN · DIAGNOSE Where is time spent?Run the checklistKeep evidence: dumps,traces, pod metrics AFTER · FIX + PREVENT Fix the code pathRedeploy via canaryAdd the missingguard to review At ~10 min: if the new version is the only change and rollback is safe, roll back first.
Mitigation comes before root cause. Capture evidence on the way so the rollback does not erase it.

Three conditions make rollback the right call: the timeouts began within minutes of the rollout, the new version is the only meaningful change in that window, and the previous version can run against the current data and schema. If a migration already dropped a column the old code reads, or the new version wrote data the old version cannot parse, rollback becomes its own incident. In that case, turning off a feature flag, scaling out, or shedding the new endpoint may be safer. The full decision is covered in How to Decide Whether to Roll Back a Release.

Before you roll back, spend sixty seconds capturing what will disappear with the bad pods:

Shell · capture, then roll back
# Which revision is live, and when did it start rolling out?
kubectl rollout history deployment/orders-api
kubectl get pods -l app=orders-api -o wide --sort-by=.metadata.creationTimestamp

# Snapshot resource use and recent logs from one new pod before it goes away
kubectl top pods -l app=orders-api
kubectl logs orders-api-7c9f8d6b5-x2x4q --since=15m > new-pod.log

# Thread / goroutine / heap dump if your runtime supports it (JVM example)
kubectl exec orders-api-7c9f8d6b5-x2x4q -- jcmd 1 Thread.print > threads.txt

# Roll back to the previous revision and watch it complete
kubectl rollout undo deployment/orders-api
kubectl rollout status deployment/orders-api
Don’t raise the timeout as a mitigation

A longer timeout does not make the slow hop faster. It lets more requests pile up behind it, each holding a thread, a connection, and memory, so the service reaches its limits sooner. Raise a timeout only when evidence shows the work is legitimately slower and the capacity to carry it exists.

Is the deployment really the cause?

A deploy is the most visible change in any incident window, so it gets blamed first. Often it deserves to be, but check before you act on that assumption. The fastest test is to compare the new version with the old one during the rollout itself, while both are serving traffic.

If you need to go further, Which Commit Caused the Production Incident? covers narrowing a release to the change that caused it, and How to Correlate Logs, Traces, and a Code Change During an Incident covers stitching the signals together. Developers hit the same ambiguity often; searches for 504 gateway timeout after deployment return many threads where the deploy was blamed for a load balancer or dependency change.

The API timeout triage checklist

Run these steps in order. Each one either narrows the cause or rules out a group of causes, and the order front-loads the checks that are cheapest and most decisive.

  1. Confirm and scope. Check that timeouts began with the rollout, and find which endpoints, pods, versions, and regions are affected.
  2. Mitigate. Roll back, disable the feature flag, or scale out, after capturing logs, dumps, and pod metrics from a new pod.
  3. Find the reporting layer and its limit. Identify which component gave up (client, gateway, service, or driver) and what its timeout value is.
  4. Read the latency distribution. Decide whether all requests got slower, a subset is pinned at a timeout value, or one group is slow.
  5. Check resources on new pods. Look at CPU throttling, memory and GC, restarts, and the number of ready replicas.
  6. Check for waiting. Look for requests queued on a connection pool, a database lock, a thread pool, or a blocked event loop.
  7. Trace one slow request. Compare a slow trace from the new version with a normal trace from the old one and find the span that grew or appeared.
  8. Diff the configuration, not just the code. Compare rendered environment variables, timeouts, endpoints, pool sizes, and probes between the two versions.
Scope
  • New pods only, or all pods?
  • One endpoint, or every endpoint?
  • One region, zone, or tenant?
  • Started with the rollout, not before?
Resources
  • CPU throttling on new pods
  • Memory, GC pauses, restarts
  • Ready replicas vs desired
  • Startup and warm-up time
Waiting
  • Connection pool wait time
  • Database lock waits
  • Thread pool or event loop lag
  • Queue depth in front of workers
Config
  • Timeout values at each hop
  • Dependency URLs and regions
  • Pool sizes and limits
  • Readiness probe and rollout strategy

Which of the five causes is it?

Four questions, asked in order, sort almost every post-deploy timeout into one of five groups. Each question is answerable from standard metrics and one or two traces.

FOUR QUESTIONS, IN ORDER Q1 · SPLIT BY VERSION Do old-version pods time out too? Q2 · POD METRICS Are CPU, memory, or throttling up on new pods only? Q3 · WAIT TIME Do requests wait on a pool, lock, or event loop? Q4 · ONE SLOW TRACE Does a slow trace show a new or slower call? Shared cause Dependency, load, migration Resource starvation CPU limit, GC, fewer pods Contention Pool, DB lock, event loop New code path Query, N+1 loop, remote call YESYESYESYES NONONONO Config drift Diff rendered timeouts, endpoints, env vars, pool sizes, and probes between the two versions.
Splitting by version first saves the most time: it tells you whether to look at the release or at something it shares.

1. Shared cause: old pods are slow too

If both versions time out, the release is probably not the cause, or it is only the trigger. Look for a schema migration that ran with the deploy, a dependency that deployed at the same time, a batch job, a traffic spike, or an infrastructure event. Migrations are the most common case that looks like a code problem: they ship with the release but act on the shared database.

2. Resource starvation on new pods

The new version needs more CPU or memory per request than the old one, or it has less to work with. Typical triggers are a heavier serialization library, a new per-request computation, larger in-memory caches, a changed CPU limit in the manifest, or a rollout that left fewer ready replicas. CPU throttling is easy to miss because CPU usage looks capped rather than high; the cAdvisor counters container_cpu_cfs_throttled_periods_total over container_cpu_cfs_periods_total show the throttled fraction.

3. Contention: requests wait behind something

Each request does roughly the same work as before, but spends much longer waiting for a shared resource: a database connection, a row or table lock, a worker thread, or the event loop in Node.js. Latency stays normal for requests that get through and jumps for those that queue. If connection pools are involved, HikariCP: Connection Is Not Available, Request Timed Out and Connection Leak vs. Pool Exhaustion go deeper on telling a leak from a long hold.

4. A new or slower code path

The release added work on a hot path: a query without an index, a loop that issues one query per item, a synchronous call to another service, or a larger response payload. A trace from the new version shows a span that was not there before, or a familiar span that is now much longer. The query side of this is covered in Why an API Is Slow in Production but Fast Locally.

5. Config drift

The code is fine but the environment around it changed: a timeout value in a new config file, a dependency URL pointing at another region, a pool size that fell back to its default when a key was renamed, or a readiness probe that now passes before the app is warm. Diff the rendered configuration of the two versions (environment, mounted config, Helm values after templating), not just the source code.

How do you read latency and timeout signals?

The shape of the latency distribution after the deploy is the single most informative signal you have. The average hides it; look at a histogram or at p50, p95, and p99 side by side.

THREE SHAPES OF A POST-DEPLOY LATENCY CHANGE before deploy after deploy timeout latency →latency →latency → Whole distribution shifts Every request does more work: CPU, GC, new per-request cost Spike at the timeout Most requests fine; a group waits until a timer fires A second hump One subset is slow: endpoint, pod, tenant, or cache miss
A shift points at work per request. A spike at a round number points at waiting. A second hump points at one subset of traffic.
Signal after the deployWhat it usually meansCheck next
p50 and p99 both roseEvery request does more work, or has less CPU to do it withCPU throttling, GC time, a new span on the common path
p50 flat, p99 pinned at a round number (5 s, 15 s, 30 s)A group of requests waits until some timer firesFind which component has that timeout value; look for pool or lock waits behind it
Timeouts only on new-version podsThe release or its configTrace diff between versions; rendered config diff
Timeouts on old and new pods alikeA shared change: migration, dependency, trafficDatabase locks, dependency deploys, request rate
Worst in the first minutes, then easingWarm-up: cold caches, JIT, connection setup, a probe passing too earlyReadiness probe, cache hit rate, startup connection count
Getting steadily worse after the rollout finishedSomething accumulates: a leak, a growing queue, a retry loopPool active count, queue depth, retry metrics, memory trend
504 with a short durationA hop with a shorter timeout than expected, or a connect failureGateway route timeouts, connect timeouts, DNS or endpoint changes
Only one endpoint or tenant affectedA code path or data shape specific to that trafficTraces for that route; data volume for that tenant
Round numbers are clues

When failed requests all take almost exactly 5, 10, 15, 30, or 60 seconds, a timer somewhere is deciding their fate. Search your config for that value. The component that owns it is the one giving up, and whatever it is waiting on is the next hop to inspect. For example, nginx’s proxy_read_timeout defaults to 60 seconds, and an Envoy route timeout defaults to 15 seconds.

Which deploy-specific causes do people miss?

Some causes exist only because a deployment happened. They are easy to miss because the code diff looks harmless.

A migration waiting on a lock

In PostgreSQL, most ALTER TABLE forms need an ACCESS EXCLUSIVE lock. If a long transaction holds any lock on the table, the ALTER waits, and every query that arrives after it queues behind the ALTER, including plain reads. A quick migration becomes minutes of timeouts on every endpoint that touches the table. CREATE INDEX without CONCURRENTLY blocks writes to the table while it builds.

PostgreSQL · who is blocking whom
SELECT pid, pg_blocking_pids(pid) AS blocked_by,
       state, wait_event_type,
       now() - query_start AS waiting_for,
       left(query, 70) AS query
FROM   pg_stat_activity
WHERE  cardinality(pg_blocking_pids(pid)) > 0
ORDER BY query_start;
migration.sqlcan queue every query behind it
ALTER TABLE orders ADD COLUMN fulfillment_region text;
CREATE INDEX orders_region_idx ON orders (fulfillment_region);  -- blocks writes
migration.sqlfails fast instead of queueing traffic
SET lock_timeout = '3s';   -- fail fast instead of queueing traffic; retry later
ALTER TABLE orders ADD COLUMN fulfillment_region text;

-- run outside a transaction block
CREATE INDEX CONCURRENTLY orders_region_idx ON orders (fulfillment_region);

The PostgreSQL documentation describes lock_timeout and statement_timeout; both are worth setting for migration sessions.

A new call without a timeout

The most common post-deploy code cause is a new synchronous call to another service on a hot path. If the client has no timeout, one slow dependency holds a worker per request until something upstream gives up. Python’s requests, for example, never times out unless you pass timeout.

orders/handlers.pyno timeout, no fallback
def get_order(order_id):
    order = repo.load(order_id)
    # added in this release: show estimated delivery on every order read
    eta = requests.get(f"{SHIPPING_URL}/eta/{order_id}").json()   # waits forever
    return {**order.to_dict(), "eta": eta}
orders/handlers.pybounded, degrades gracefully
def get_order(order_id):
    order = repo.load(order_id)
    try:
        resp = shipping.get(f"{SHIPPING_URL}/eta/{order_id}",
                            timeout=(0.2, 0.5))   # connect, read seconds
        resp.raise_for_status()
        eta = resp.json()
    except requests.RequestException:
        eta = None   # ETA is optional; the order is not
    return {**order.to_dict(), "eta": eta}

Capacity lost during the rollout

A rolling update removes old pods as new ones become ready. With the Kubernetes defaults of maxUnavailable: 25% and maxSurge: 25%, a service already running near its limit briefly has less capacity than it needs. If readiness probes pass before the application has warmed its caches, JIT, or connection pools, the new pods also start slower than the old ones they replace. See the Kubernetes Deployment documentation for the rolling update settings.

Cold caches and connection storms

A release that changes a cache key format, a serialization version, or a cache namespace starts with an empty cache, so every request goes to the database at once. That is the pattern in Cache Stampede: Why a Healthy Cache Can Overload Your Database. Similarly, many new pods opening full connection pools at the same moment can push a database to its connection limit.

The component that reports the timeout is the one that gave up. The fix belongs to the one that made it wait.

Mitigation and durable fix, by cause

CauseMitigate nowDurable fix
New code pathRoll back or disable the flag for that pathFix the query or loop, add a timeout and fallback to the new call, load test the route
Resource starvationScale out, or restore the previous resource limitsProfile the new per-request cost; set limits from measured usage
Contention (pool, lock, event loop)Roll back; kill the blocking session if it is a migrationShorter holds, no I/O inside transactions, lock_timeout on migrations, move CPU work off the event loop
Shared causePause the job or migration; coordinate with the dependency ownerSeparate migrations from releases; expand-and-contract schema changes
Config driftRestore the previous config valuesReview rendered config in the pull request; validate required keys at startup
Warm-up and rollout capacityPause the rollout; scale up before resumingReadiness that reflects warm state; lower maxUnavailable; staged or canary rollout

After the service is stable, ask what the release checks missed. A new call without a timeout, a migration without a lock timeout, or a resource limit change are all visible in the pull request before merge; they are rarely visible in unit tests. A PR Passed CI but Broke Production looks at that gap directly, and the Google SRE book chapter on managing incidents covers the coordination side of the first thirty minutes.

How Tomosu helps

Tomosu analyzes a repository and each pull request for production reliability risk. For post-deploy timeouts, the value is in seeing, before the release, the risks this checklist looks for after it:

These findings feed the Production Reliability Index, so a release that adds a slow dependency to a hot path is visible to reviewers before it ships. To try it on your own code, follow the repository scan guide.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

Why do API timeouts start right after a deployment?

The most common reasons are a new or slower code path on a hot endpoint, less CPU or memory per request on the new pods, requests waiting on a connection pool or database lock, a shared change such as a migration that shipped with the release, or configuration drift such as changed timeouts or endpoints. Warm-up and reduced capacity during a rolling update can also cause short bursts of timeouts.

Should I roll back or debug first when timeouts follow a deploy?

Roll back first if the timeouts began with the rollout, the new version is the only meaningful change, and the previous version can run against the current schema and data. Capture logs, dumps, and pod metrics from a new pod before rolling back so you can diagnose the bad version afterwards.

Why does the gateway return 504 when the problem is in my service?

A 504 Gateway Timeout means the gateway or load balancer stopped waiting for your service. The gateway reports the error because its timeout was the shortest one on the path. The time was spent inside your service or one of its dependencies, so start from the gateway and trace inward.

Should I increase the timeout to stop the errors?

Not as a first response. A longer timeout lets more requests pile up behind the slow component, each holding a thread, connection, and memory, which usually makes the outage wider. Increase a timeout only when evidence shows the work is legitimately slower and there is capacity to carry the extra in-flight requests.

How can I tell if the deployment or something else caused the timeouts?

Split latency and errors by version during the rollout. If only new-version pods time out, the release or its configuration is implicated. If old pods time out too, look for a shared change in the same window, such as a database migration, a dependency deploy, a batch job, or a traffic spike.

Can a database migration cause API timeouts after a deploy?

Yes. In PostgreSQL, most ALTER TABLE forms need an ACCESS EXCLUSIVE lock. If the migration waits behind a long transaction, later queries on that table queue behind it, including reads. Setting lock_timeout on the migration session makes it fail fast instead of blocking traffic, and CREATE INDEX CONCURRENTLY avoids blocking writes while an index builds.

What does it mean when failed requests all take exactly 30 seconds?

A timer somewhere on the request path is set to 30 seconds and is deciding when those requests fail. Search your gateway, client, driver, and pool configuration for that value. The component that owns it is the one giving up, and the thing it is waiting on is the next hop to inspect.


Most post-deploy timeouts were visible in the pull request: a new call with no timeout, a migration with no lock timeout, a limit that changed. Tomosu surfaces those before the release. Assess your repository →