Most root cause analysis documents answer one of three questions well. Some name the trigger (“a flag was turned on”). Some describe the symptom (“the cart returned 504s”). Some list what users saw. Very few connect them, so readers cannot tell why that trigger produced that impact, and the action items end up aimed at the wrong link.
A useful RCA is a causal chain with evidence at every link: the trigger (what changed), the code path (which code turned that change into a failure), the failure mode (how the system broke), and the user impact (who saw what, for how long). Add the conditions that had to be true and why the chain was not caught earlier.
- Trigger: the event that started it, with a timestamp and source.
- Code path: entry point to failing call, with file, function, and commit.
- User impact: measured, bounded, and stated in user terms.
- Action items: each one breaks a named link in the chain.
This guide shows how to build that chain, how to write each section, and how to review an RCA before it is published. It uses one worked example throughout: a hypothetical release where turning on a feature flag made the cart page fail. The example is invented, but the pattern is common. The broader culture behind it, blameless postmortems written to change the system rather than assign fault, is described in Google’s SRE book chapter on postmortem culture.
What should an RCA explain?
An RCA (root cause analysis) is the section of an incident review that explains why the incident happened, in enough detail that someone who was not on the call could predict it from the facts. A timeline says what happened in order. An RCA says what caused what.
For software incidents, that explanation has four links. If any one is missing, the reader has to guess, and the action items usually target whatever link the author understood best.
- Trigger: the event that started the incident. A deploy, a flag change, a config push, a traffic spike, a certificate expiry, a dependency outage.
- Code path: the specific code that turned the trigger into a failure, from the entry point (an endpoint, a consumer, a cron job) down to the call that failed or blocked.
- Failure mode: how the system broke in operational terms. Thread or worker exhaustion, a crash loop, a lock pile-up, a queue backlog, corrupt writes.
- User impact: what users experienced, how many, for how long, and whether anything was lost or has to be repaired.
The chain matters because each link suggests a different kind of fix. Preventing the trigger is often impossible (traffic spikes will happen). Hardening the code path usually has the most leverage. Containing the failure mode limits the blast radius. Detecting impact sooner shortens the next incident. An RCA that names only one link can only produce action items for that link.
What is the difference between a trigger and a root cause?
The trigger is the event that started the incident; the root cause is the latent weakness in the system that let that event cause harm. In the example, the trigger is a flag rollout. The root cause is that the cart endpoint made a synchronous call to a non-critical service with no fallback and a timeout longer than its caller’s. The same trigger on a well-built code path would have caused a few missing recommendations, not a failed cart.
| Term | Definition | In the worked example |
|---|---|---|
| Trigger | The event that started the incident | cart_recommendations flag raised from 0% to 100% at 14:02 |
| Root cause | The latent weakness in code or design that the trigger exposed | Blocking recs call on the request path with no fallback and a 10 s timeout |
| Contributing factor | A condition that made the incident possible or worse | Campaign traffic, 8 sync workers per pod, gateway timeout (5 s) shorter than client timeout |
| Failure mode | How the system broke, operationally | All workers blocked on recs; new requests queued until the gateway gave up |
| Detection gap | Why it took as long as it did to notice | Alerting on checkout conversion only; no alert on cart latency or 5xx |
| Escape path | Why tests, review, and rollout did not stop it | Tests mocked the recs client; flag went 0→100% in one step |
Complex incidents rarely have a single root cause. Google’s example postmortem in the SRE book lists root causes and the trigger under separate headings, and the SRE workbook’s postmortem chapter walks through what separates a useful postmortem from a weak one. A practical rule: if removing a condition would have prevented the incident, list it. The chain format makes that natural, because each link can have more than one supporting condition.
Why do most RCAs fail to connect the trigger, code path, and impact?
Most weak RCAs are not wrong. They stop one link early, or they replace a link with a label. Here are the patterns that show up most often, and what each one is missing.
| What the RCA says | What is missing | What to write instead |
|---|---|---|
| “Root cause: bad deploy.” | The code path. Which change, which function, which behavior? | The commit, the function it changed, and the behavior that changed under what condition |
| “The database was slow.” | The cause of the symptom. Slow because of what, and why did slowness break users? | The query or lock, what made it slow, and which callers had no bound on waiting for it |
| “Human error: engineer enabled the flag.” | The system condition. Why could one routine action take down the cart? | Why the rollout had no stages or automatic check, and why the code had no fallback |
| A 60-line timeline, no analysis | Causation. Order is not cause. | A short causal chain above the timeline, with timeline entries as its evidence |
| “Some users saw errors.” | Scope and duration. Nobody can prioritize the fix. | Who, what they saw, what share, how long, and any data to repair |
| “Action: add more tests.” | A target. Which test would have failed, and on which link? | “Add a test where the recs client times out and assert the cart still returns 200” |
The common thread is that the document describes the incident from the outside. Operators see triggers and symptoms. Users see impact. The code path is the part only engineers can supply, and it is the part most often skipped because it takes the most work to establish.
A timeline tells you what happened in order. An RCA tells you why the first event could cause the last one.
How do you write an RCA, step by step?
Write the RCA in this order, even though it will be read in a different one. Starting from impact keeps the analysis honest about what actually mattered; ending with action items keeps them tied to evidence.
- Pin down the user impact first. Who was affected, what they experienced, how many, from when to when, and whether data needs repair. Use measured numbers and say where they came from.
- Build the timeline from signals, not memory. Pull timestamps from deploy logs, flag audit logs, alerts, dashboards, and chat. Mark the first user-visible failure and the moment it ended.
- Name the trigger with its source. One event, a timestamp, and the record that proves it. If the code shipped earlier than the trigger, say so; the commit and the trigger are often different events.
- Trace the code path. From the entry point the users hit to the call that failed or blocked, naming files, functions, and the commit that introduced the behavior.
- Explain why it failed now and not before. List the conditions (traffic, data shape, config, dependency state) that had to be true. This is where contributing factors come from.
- Explain why it was not caught earlier. Which test, review step, rollout stage, or alert could have stopped or shortened it, and why it did not.
- Write action items against links in the chain. Each item names the link it breaks, an owner, and how you will know it is done.
Steps 2 and 3 are where most of the investigation happens during the incident itself. If you still need to establish which change was running, Which Commit Caused the Production Incident? covers deploy-to-commit mapping and bisecting, and How to Correlate Logs, Traces, and a Code Change During an Incident covers stamping the git SHA into traces and logs so the evidence exists when you need it. This post assumes you have that evidence and focuses on writing it up.
How do you trace and write the code path section?
The code path section explains how the trigger became a failure inside your code. It should let a reader open the repository and follow the path without asking anyone. That means naming the entry point, each hop, and the line that behaved badly, plus the commit that introduced it.
Start from the evidence closest to the user and walk inward. In the example, a trace of a failed GET /cart shows the gateway span ending at 5 s with a 504, while the application span underneath keeps running and contains a recs.fetch child span of 9.8 s. That child span points to the client, and git blame on the call site points to the pull request that added it.
# cart/service.py
def build_cart_response(cart, user):
body = serialize_cart(cart)
if flags.enabled("cart_recommendations", user):
# Blocking call on the request worker. Any exception fails the whole cart.
body["recommendations"] = recs_client.fetch(
user_id=user.id, sku_ids=[i.sku for i in cart.items],
)
return body
# recs/client.py
class RecsClient:
def fetch(self, user_id, sku_ids):
resp = self.session.get(
f"{RECS_URL}/v1/recommendations",
params={"user": user_id, "skus": ",".join(sku_ids)},
timeout=10, # longer than the gateway's 5 s route timeout
)
resp.raise_for_status()
return resp.json()["items"]
Nothing in that diff is obviously wrong in isolation, which is why it passed review. The failure only makes sense once you add two facts from outside the diff: the service runs a small number of synchronous workers per pod, and the gateway gives up after 5 seconds. Put those facts in the RCA, with where they come from.
A good code path section in the written RCA reads something like this:
GET /cart is served by CartView.get, which calls build_cart_response (cart/service.py). PR #2217, deployed in v2.41 at 13:41, added a call to RecsClient.fetch inside that function when the cart_recommendations flag is on. The call is synchronous, has a 10 s timeout, and any exception propagates and fails the request.
Each pod runs 8 synchronous workers. When the recommendations service slowed to a p99 of about 8 s under campaign traffic, each cart request held a worker for that long. The gateway’s 5 s route timeout returned 504 to users while the workers kept waiting, so pods spent capacity on abandoned requests and queued new ones. Evidence: trace 4bf92f35… (gateway span 5.0 s, recs.fetch span 9.8 s), worker busy metric at 8/8 on all pods from 14:05.
Two details make this section credible. It separates the introducing change (PR #2217, deployed at 13:41) from the trigger (the flag at 14:02), and it cites the evidence for each claim. It also avoids blame: it says what the code did, not who wrote it.
Show the fix in the same terms
If the remediation is a code change, include it or link it, and explain which link it breaks. For the example, the fix bounds the call well inside the gateway’s budget and degrades to an empty list instead of failing the cart:
import requests
RECS_TIMEOUT = (0.1, 0.3) # (connect, read) seconds, far inside the 5 s gateway budget
def build_cart_response(cart, user):
body = serialize_cart(cart)
body["recommendations"] = []
if flags.enabled("cart_recommendations", user):
try:
body["recommendations"] = recs_client.fetch(
user_id=user.id, sku_ids=[i.sku for i in cart.items],
timeout=RECS_TIMEOUT,
)
except requests.RequestException:
# Recommendations are optional. Degrade; never fail the cart.
metrics.incr("cart.recs.fallback")
return body
One caveat belongs in the RCA if you use Python requests: its read timeout limits the wait between bytes from the server, not the total response time. For a hard upper bound on the whole call, add an overall deadline or a circuit breaker around the client. The choice between those controls is covered in Circuit Breaker vs. Retry vs. Load Shedding.
How do you measure and state user impact in an RCA?
State user impact as who was affected, what they experienced, how many, for how long, and what has to be repaired, each with its data source. “Degraded performance for some users” cannot be prioritized or compared with the next incident. A specific statement can.
| Dimension | How to measure it | Example statement |
|---|---|---|
| Who | Segment by region, platform, plan, tenant, or flag cohort | All signed-in web and app users; guest carts unaffected (flag was user-targeted) |
| What they saw | Status codes at the edge, client error events, support tickets | Cart page failed to load with a generic error; checkout could not be started |
| How many | Edge logs or RUM, counted in sessions or users, not only requests | 31% of cart page views; about 18,400 sessions saw at least one failure |
| How long | First and last user-visible failure, not alert times | 14:06 to 14:29 (23 minutes) |
| Data consequences | Reconciliation queries, audit logs | No orders lost or duplicated; no data repair needed |
| Downstream effects | Business metrics, partner SLAs, error budget | Checkout starts below forecast for the window; cart SLO budget for the month consumed |
Client retries, page reloads, and polling inflate request-level error counts, and saturated services can under-count failures because requests never reach the application logs. Measure at the edge or in real user monitoring where you can, count sessions or users as well as requests, and say which measure you used. If a number is an estimate, label it as one and give the method.
Also state what did not happen. “No data loss” and “payments unaffected” are findings too, and they need the same evidence. They are often the first thing a stakeholder asks.
What does a good RCA template look like?
This template puts the causal chain near the top, where readers look first, and keeps the timeline as supporting evidence. Copy it into your incident tool or docs system and keep the headings stable, so RCAs can be compared across incidents.
# RCA: <one-line summary: trigger → failure → impact>
Status: draft | in review | final Severity: SEV-n
Incident window: <first user-visible failure> – <last> (<duration>)
## 1. User impact
Who / what they experienced / how many (source) / how long
Data to repair / what was NOT affected
## 2. Causal chain
Trigger: <event, timestamp, source record>
Code path: <entry point> → <file:function> → <failing call> (introduced in <PR/commit>)
Failure mode: <how the system broke: exhaustion, crash loop, backlog, bad writes>
User impact: <one line, links to section 1>
## 3. Contributing factors
- <condition that had to be true> — evidence: <link>
## 4. Why it was not caught earlier
Tests: ... Review: ... Rollout: ... Detection: ...
## 5. Timeline (evidence for the chain)
HH:MM <event> <source>
## 6. What went well / where we got lucky
## 7. Action items
| Action | Link it breaks | Owner | Due | Done when |
## 8. Open questions
A few rules make the template work in practice:
- The summary line is the chain in one sentence. “Enabling cart recommendations at 100% during campaign traffic blocked all cart workers on a slow recommendations call, failing 31% of cart views for 23 minutes.”
- Every claim has a source. A dashboard link, a trace id, a log query, a commit. “We believe” is allowed, but it goes in open questions.
- “Where we got lucky” is required. If the incident would have been worse at a different hour or with different traffic, that is a finding.
- No names in the causal chain. Describe actions and systems. People appear in the timeline only as roles.
How do you turn an RCA into action items that break the chain?
Each action item should name the link in the causal chain it breaks and how you will verify it is done. That rule removes most weak action items on its own, because “be more careful” and “improve monitoring” do not break any specific link.
| Action item | Link it breaks | Done when |
|---|---|---|
| Bound recs call to 300 ms read timeout, fall back to empty list | Code path | Merged, and a test with a stalled recs stub asserts 200 within the gateway budget |
| Cap concurrent recs calls per pod | Failure mode | Load test with recs at 8 s p99 keeps cart p99 under target |
| Audit client timeouts against gateway route timeouts | Failure mode | No client timeout on a user request path exceeds its route timeout |
| Staged flag rollout with automatic halt | Trigger | Flag service enforces stages for flags tagged request-path |
| Alert on cart 5xx and latency SLO burn | Detection | Alert fires in a replay of this incident’s metrics within 3 minutes |
Two more checks help. First, ask whether the action items would have prevented a different trigger through the same code path, such as a recommendations deploy instead of a flag. If yes, you have hardened the path rather than patched the event. Second, look for the same pattern elsewhere. The cart was one endpoint; other endpoints may make the same kind of blocking call to optional services. That search is often the most valuable action item in the document. For a related discussion of why these patterns pass tests, see A PR Passed CI but Broke Production.
How do you review an RCA before publishing it?
Review an RCA the way you would review code: against a checklist, by someone who was not the author. The questions below catch most gaps.
- Trigger has a timestamp and a source record
- Code path names entry point, function, and commit
- Failure mode explains why the path broke then
- Impact is measured, bounded, and sourced
- Every claim links to a trace, log, metric, or diff
- Hypotheses are labeled and in open questions
- Impact counts say whether they are requests, sessions, or users
- Each action names the link it breaks
- Each has an owner, a due date, and a done-when
- At least one hardens the code path, not just the trigger
- Someone searched for the same pattern elsewhere
If the incident involved an AI coding agent or an automated change, the same chain applies with one more link to document: the agent’s actions and tool calls that produced the change. The postmortem template for incidents an AI agent caused covers that extra section.
How Tomosu helps
The hardest section of an RCA to write is the code path, because it needs context that is spread across the repository: which endpoints reach a function, what timeouts and retries sit around a call, and which change introduced the behavior. Tomosu analyzes the codebase and each change for production reliability risk, which gives you that context before and after an incident:
- Call paths and dependencies: which entry points reach a changed function, and which remote calls, transactions, and shared resources sit on that path.
- Risky patterns on request paths: blocking calls to other services without a bound or fallback, timeouts that exceed their caller’s, and retries that stack across layers.
- Blast radius of a change: which services and user flows depend on the code a pull request touches, which is the starting point for the impact section.
- The same pattern elsewhere: once an RCA names a risky pattern, a repository scan shows where else it occurs, so the “search for siblings” action item has a concrete list.
These signals roll up into the Production Reliability Index. The goal is modest: fewer incidents where the code path section is the first time anyone traced the path.
Scan your repository with Tomosu →
Key takeaways
- An RCA is a causal chain, not a timeline: trigger, code path, failure mode, and user impact, with evidence at each link.
- The trigger is the event; the root cause is the weakness the event exposed. The introducing commit and the trigger are often different events.
- The code path section should let a reader follow the failure in the repository: entry point, functions, failing call, and the change that introduced it.
- State impact as who, what, how many, how long, and what needs repair, and say whether counts are requests, sessions, or users.
- Contributing factors are the conditions that had to be true. List each one with its evidence.
- Every action item should name the link it breaks and how you will verify it. Harden the code path, not only the trigger.
- Search for the same pattern elsewhere. It is often the most valuable action item.
Frequently asked questions
What is the difference between a trigger and a root cause in an RCA?
The trigger is the event that started the incident, such as a deploy, a feature flag change, a config push, or a traffic spike. The root cause is the latent weakness the trigger exposed, such as a blocking call with no fallback or a timeout longer than its caller’s. The same trigger on a well-built code path would cause little or no harm, which is why an RCA should name both.
What should an RCA include?
A one-line summary of the chain, a measured user impact statement, the causal chain (trigger, code path, failure mode, user impact), contributing factors with evidence, why tests, review, rollout, and alerting did not catch it, a timeline as supporting evidence, what went well and where you got lucky, action items tied to links in the chain, and open questions.
How do you describe user impact in an RCA?
Say who was affected, what they experienced, how many were affected, from when to when, and whether any data must be repaired, with the source for each number. Measure at the edge or with real user monitoring where possible, and state whether counts are requests, sessions, or users, because retries and reloads inflate request counts.
Is the 5 whys method enough for a software RCA?
It is a useful prompt, but on its own it tends to produce a single linear story that stops at whichever answer the group agrees on, often a person. Software incidents usually have several contributing conditions. Use the questions to explore, then write the result as a causal chain with evidence and a list of contributing factors.
How detailed should the code path section of an RCA be?
Detailed enough that an engineer who was not involved can follow it in the repository: the entry point users hit, each function or service hop, the call that failed or blocked, the commit or pull request that introduced the behavior, and the facts from outside the diff, such as worker counts or gateway timeouts, that made it fail.
What makes a good RCA action item?
A good action item names the link in the causal chain it breaks, has a single owner and a due date, and has a verifiable done-when condition, such as a test that fails before the fix and passes after. Items like “be more careful” or “improve monitoring” break no specific link and rarely change anything.
Who should write the RCA, and how soon?
The people closest to the incident, usually the incident commander and the engineers who owned the affected code, should draft it while evidence and memory are fresh, typically within a few working days. Someone who was not involved should review it. Keep it blameless: describe what systems and processes allowed, not who made a mistake.
The best RCAs read like a proof: this event, through this code, caused this harm, and these changes break the chain. Tomosu maps the code paths so that proof is easier to write, and so fewer incidents need one. Assess your repository →