Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Incident response

How to Write an RCA That Connects the Trigger, Code Path, and User Impact

Tomosu AI·15 min read·

Most root cause analysis documents answer one of three questions well. Some name the trigger (“a flag was turned on”). Some describe the symptom (“the cart returned 504s”). Some list what users saw. Very few connect them, so readers cannot tell why that trigger produced that impact, and the action items end up aimed at the wrong link.

Quick answer

A useful RCA is a causal chain with evidence at every link: the trigger (what changed), the code path (which code turned that change into a failure), the failure mode (how the system broke), and the user impact (who saw what, for how long). Add the conditions that had to be true and why the chain was not caught earlier.

This guide shows how to build that chain, how to write each section, and how to review an RCA before it is published. It uses one worked example throughout: a hypothetical release where turning on a feature flag made the cart page fail. The example is invented, but the pattern is common. The broader culture behind it, blameless postmortems written to change the system rather than assign fault, is described in Google’s SRE book chapter on postmortem culture.

What should an RCA explain?

An RCA (root cause analysis) is the section of an incident review that explains why the incident happened, in enough detail that someone who was not on the call could predict it from the facts. A timeline says what happened in order. An RCA says what caused what.

For software incidents, that explanation has four links. If any one is missing, the reader has to guess, and the action items usually target whatever link the author understood best.

AN RCA IS A CAUSAL CHAIN, NOT A TIMELINE 1 · TRIGGER Flag set to 100% What changed, and when: 14:02 2 · CODE PATH Sync recs call in GET /cart, with a 10 s timeout 3 · FAILURE MODE Workers exhausted Requests queue, gateway returns 504 4 · USER IMPACT Cart fails to load 31% of cart views for 23 minutes EVIDENCE FOR EACH LINK flag audit log trace spans + SHA worker metrics RUM + edge logs CONDITIONS THAT HAD TO BE TRUE (CONTRIBUTING FACTORS) Campaign traffic slowed the recommendations service · no fallback when recs time out 8 synchronous workers per pod · client timeout longer than gateway timeout · no cart p99 alert
Each link is a claim, and each claim has its own evidence. Contributing factors are the conditions that let the trigger travel all the way to users.

The chain matters because each link suggests a different kind of fix. Preventing the trigger is often impossible (traffic spikes will happen). Hardening the code path usually has the most leverage. Containing the failure mode limits the blast radius. Detecting impact sooner shortens the next incident. An RCA that names only one link can only produce action items for that link.

What is the difference between a trigger and a root cause?

The trigger is the event that started the incident; the root cause is the latent weakness in the system that let that event cause harm. In the example, the trigger is a flag rollout. The root cause is that the cart endpoint made a synchronous call to a non-critical service with no fallback and a timeout longer than its caller’s. The same trigger on a well-built code path would have caused a few missing recommendations, not a failed cart.

TermDefinitionIn the worked example
TriggerThe event that started the incidentcart_recommendations flag raised from 0% to 100% at 14:02
Root causeThe latent weakness in code or design that the trigger exposedBlocking recs call on the request path with no fallback and a 10 s timeout
Contributing factorA condition that made the incident possible or worseCampaign traffic, 8 sync workers per pod, gateway timeout (5 s) shorter than client timeout
Failure modeHow the system broke, operationallyAll workers blocked on recs; new requests queued until the gateway gave up
Detection gapWhy it took as long as it did to noticeAlerting on checkout conversion only; no alert on cart latency or 5xx
Escape pathWhy tests, review, and rollout did not stop itTests mocked the recs client; flag went 0→100% in one step
“Root cause” is usually plural

Complex incidents rarely have a single root cause. Google’s example postmortem in the SRE book lists root causes and the trigger under separate headings, and the SRE workbook’s postmortem chapter walks through what separates a useful postmortem from a weak one. A practical rule: if removing a condition would have prevented the incident, list it. The chain format makes that natural, because each link can have more than one supporting condition.

Why do most RCAs fail to connect the trigger, code path, and impact?

Most weak RCAs are not wrong. They stop one link early, or they replace a link with a label. Here are the patterns that show up most often, and what each one is missing.

What the RCA saysWhat is missingWhat to write instead
“Root cause: bad deploy.”The code path. Which change, which function, which behavior?The commit, the function it changed, and the behavior that changed under what condition
“The database was slow.”The cause of the symptom. Slow because of what, and why did slowness break users?The query or lock, what made it slow, and which callers had no bound on waiting for it
“Human error: engineer enabled the flag.”The system condition. Why could one routine action take down the cart?Why the rollout had no stages or automatic check, and why the code had no fallback
A 60-line timeline, no analysisCausation. Order is not cause.A short causal chain above the timeline, with timeline entries as its evidence
“Some users saw errors.”Scope and duration. Nobody can prioritize the fix.Who, what they saw, what share, how long, and any data to repair
“Action: add more tests.”A target. Which test would have failed, and on which link?“Add a test where the recs client times out and assert the cart still returns 200”

The common thread is that the document describes the incident from the outside. Operators see triggers and symptoms. Users see impact. The code path is the part only engineers can supply, and it is the part most often skipped because it takes the most work to establish.

A timeline tells you what happened in order. An RCA tells you why the first event could cause the last one.

How do you write an RCA, step by step?

Write the RCA in this order, even though it will be read in a different one. Starting from impact keeps the analysis honest about what actually mattered; ending with action items keeps them tied to evidence.

  1. Pin down the user impact first. Who was affected, what they experienced, how many, from when to when, and whether data needs repair. Use measured numbers and say where they came from.
  2. Build the timeline from signals, not memory. Pull timestamps from deploy logs, flag audit logs, alerts, dashboards, and chat. Mark the first user-visible failure and the moment it ended.
  3. Name the trigger with its source. One event, a timestamp, and the record that proves it. If the code shipped earlier than the trigger, say so; the commit and the trigger are often different events.
  4. Trace the code path. From the entry point the users hit to the call that failed or blocked, naming files, functions, and the commit that introduced the behavior.
  5. Explain why it failed now and not before. List the conditions (traffic, data shape, config, dependency state) that had to be true. This is where contributing factors come from.
  6. Explain why it was not caught earlier. Which test, review step, rollout stage, or alert could have stopped or shortened it, and why it did not.
  7. Write action items against links in the chain. Each item names the link it breaks, an owner, and how you will know it is done.
THE COMMIT IS NOT THE TRIGGER: TIMELINE OF THE EXAMPLE USER IMPACT 13:4013:5014:1014:40 13:41 deploy v2.41 code ships dark (flag off) 14:19 incident declared commander assigned 14:02 flag to 100% TRIGGER 14:27 flag turned off mitigation 14:00 campaign email traffic rises (condition) 14:15 alert fires checkout conversion drop 14:06 first cart 504s impact starts 14:29 errors normal impact ends Deploy → trigger 21 min · time to detect 9 min · detect → mitigate 12 min · impact 23 min
The deploy happened 21 minutes before anything broke. An RCA that starts from “which deploy?” alone would look for the trigger in the wrong place.

Steps 2 and 3 are where most of the investigation happens during the incident itself. If you still need to establish which change was running, Which Commit Caused the Production Incident? covers deploy-to-commit mapping and bisecting, and How to Correlate Logs, Traces, and a Code Change During an Incident covers stamping the git SHA into traces and logs so the evidence exists when you need it. This post assumes you have that evidence and focuses on writing it up.

How do you trace and write the code path section?

The code path section explains how the trigger became a failure inside your code. It should let a reader open the repository and follow the path without asking anyone. That means naming the entry point, each hop, and the line that behaved badly, plus the commit that introduced it.

Start from the evidence closest to the user and walk inward. In the example, a trace of a failed GET /cart shows the gateway span ending at 5 s with a 504, while the application span underneath keeps running and contains a recs.fetch child span of 9.8 s. That child span points to the client, and git blame on the call site points to the pull request that added it.

cart/service.py · recs/client.py · introduced in PR #2217the code path
# cart/service.py
def build_cart_response(cart, user):
    body = serialize_cart(cart)
    if flags.enabled("cart_recommendations", user):
        # Blocking call on the request worker. Any exception fails the whole cart.
        body["recommendations"] = recs_client.fetch(
            user_id=user.id, sku_ids=[i.sku for i in cart.items],
        )
    return body

# recs/client.py
class RecsClient:
    def fetch(self, user_id, sku_ids):
        resp = self.session.get(
            f"{RECS_URL}/v1/recommendations",
            params={"user": user_id, "skus": ",".join(sku_ids)},
            timeout=10,  # longer than the gateway's 5 s route timeout
        )
        resp.raise_for_status()
        return resp.json()["items"]

Nothing in that diff is obviously wrong in isolation, which is why it passed review. The failure only makes sense once you add two facts from outside the diff: the service runs a small number of synchronous workers per pod, and the gateway gives up after 5 seconds. Put those facts in the RCA, with where they come from.

THE CODE PATH, FROM ENTRY POINT TO BLOCKING CALL GET /cart gateway route timeout: 5 s cart/views.py CartView.get() cart/service.py build_cart_response() recs/client.py RecsClient.fetch(timeout=10) recommendations-service p99 about 8 s under campaign traffic PR #2217 span 9.8 s MECHANISM 1 · WORKER EXHAUSTION 8 synchronous workers per pod. Each blocked cart request holds one worker for up to 10 s, so a pod serves under 1 cart request per second. Everything else on the pod queues too. MECHANISM 2 · TIMEOUT INVERSION Gateway returns 504 at 5 s. The worker keeps waiting up to 10 s, spending capacity on a response that nobody will receive. Where the fix has most leverage The call site: bound it, and degrade.
The code path section names each hop and the facts from outside the diff (worker count, gateway timeout) that make the failure inevitable once the trigger arrives.

A good code path section in the written RCA reads something like this:

Example: code path section

GET /cart is served by CartView.get, which calls build_cart_response (cart/service.py). PR #2217, deployed in v2.41 at 13:41, added a call to RecsClient.fetch inside that function when the cart_recommendations flag is on. The call is synchronous, has a 10 s timeout, and any exception propagates and fails the request.

Each pod runs 8 synchronous workers. When the recommendations service slowed to a p99 of about 8 s under campaign traffic, each cart request held a worker for that long. The gateway’s 5 s route timeout returned 504 to users while the workers kept waiting, so pods spent capacity on abandoned requests and queued new ones. Evidence: trace 4bf92f35… (gateway span 5.0 s, recs.fetch span 9.8 s), worker busy metric at 8/8 on all pods from 14:05.

Two details make this section credible. It separates the introducing change (PR #2217, deployed at 13:41) from the trigger (the flag at 14:02), and it cites the evidence for each claim. It also avoids blame: it says what the code did, not who wrote it.

Show the fix in the same terms

If the remediation is a code change, include it or link it, and explain which link it breaks. For the example, the fix bounds the call well inside the gateway’s budget and degrades to an empty list instead of failing the cart:

cart/service.py · remediationbreaks link 2
import requests

RECS_TIMEOUT = (0.1, 0.3)  # (connect, read) seconds, far inside the 5 s gateway budget

def build_cart_response(cart, user):
    body = serialize_cart(cart)
    body["recommendations"] = []
    if flags.enabled("cart_recommendations", user):
        try:
            body["recommendations"] = recs_client.fetch(
                user_id=user.id, sku_ids=[i.sku for i in cart.items],
                timeout=RECS_TIMEOUT,
            )
        except requests.RequestException:
            # Recommendations are optional. Degrade; never fail the cart.
            metrics.incr("cart.recs.fallback")
    return body

One caveat belongs in the RCA if you use Python requests: its read timeout limits the wait between bytes from the server, not the total response time. For a hard upper bound on the whole call, add an overall deadline or a circuit breaker around the client. The choice between those controls is covered in Circuit Breaker vs. Retry vs. Load Shedding.

How do you measure and state user impact in an RCA?

State user impact as who was affected, what they experienced, how many, for how long, and what has to be repaired, each with its data source. “Degraded performance for some users” cannot be prioritized or compared with the next incident. A specific statement can.

DimensionHow to measure itExample statement
WhoSegment by region, platform, plan, tenant, or flag cohortAll signed-in web and app users; guest carts unaffected (flag was user-targeted)
What they sawStatus codes at the edge, client error events, support ticketsCart page failed to load with a generic error; checkout could not be started
How manyEdge logs or RUM, counted in sessions or users, not only requests31% of cart page views; about 18,400 sessions saw at least one failure
How longFirst and last user-visible failure, not alert times14:06 to 14:29 (23 minutes)
Data consequencesReconciliation queries, audit logsNo orders lost or duplicated; no data repair needed
Downstream effectsBusiness metrics, partner SLAs, error budgetCheckout starts below forecast for the window; cart SLO budget for the month consumed
Requests are not users

Client retries, page reloads, and polling inflate request-level error counts, and saturated services can under-count failures because requests never reach the application logs. Measure at the edge or in real user monitoring where you can, count sessions or users as well as requests, and say which measure you used. If a number is an estimate, label it as one and give the method.

Also state what did not happen. “No data loss” and “payments unaffected” are findings too, and they need the same evidence. They are often the first thing a stakeholder asks.

What does a good RCA template look like?

This template puts the causal chain near the top, where readers look first, and keeps the timeline as supporting evidence. Copy it into your incident tool or docs system and keep the headings stable, so RCAs can be compared across incidents.

rca-template.mdcopy this
# RCA: <one-line summary: trigger → failure → impact>
Status: draft | in review | final      Severity: SEV-n
Incident window: <first user-visible failure> – <last> (<duration>)

## 1. User impact
Who / what they experienced / how many (source) / how long
Data to repair / what was NOT affected

## 2. Causal chain
Trigger:       <event, timestamp, source record>
Code path:     <entry point> → <file:function> → <failing call>  (introduced in <PR/commit>)
Failure mode:  <how the system broke: exhaustion, crash loop, backlog, bad writes>
User impact:   <one line, links to section 1>

## 3. Contributing factors
- <condition that had to be true> — evidence: <link>

## 4. Why it was not caught earlier
Tests: ...   Review: ...   Rollout: ...   Detection: ...

## 5. Timeline (evidence for the chain)
HH:MM  <event>  <source>

## 6. What went well / where we got lucky

## 7. Action items
| Action | Link it breaks | Owner | Due | Done when |

## 8. Open questions

A few rules make the template work in practice:

How do you turn an RCA into action items that break the chain?

Each action item should name the link in the causal chain it breaks and how you will verify it is done. That rule removes most weak action items on its own, because “be more careful” and “improve monitoring” do not break any specific link.

EVERY ACTION ITEM BREAKS A NAMED LINK TRIGGER CODE PATH FAILURE MODE IMPACT + DETECTION Stage flag rollouts: 1%, 10%, 50%, 100% with an automatic health check that halts the ramp 300 ms read timeout and empty fallback. Test: recs times out, cart still returns 200 Cap concurrent recs calls per pod. Keep every client timeout below the gateway timeout Alert on cart 5xx and latency SLO burn. Add a synthetic cart check PREVENTHARDENCONTAINDETECT SOONER ACTION ITEMS THAT BREAK NO LINK “Be more careful with flags” · “Improve monitoring” · “Add more tests” · “Review PRs more closely”
Spread action items across links. Hardening the code path usually gives the most protection, because it holds no matter what triggers the next incident.
Action itemLink it breaksDone when
Bound recs call to 300 ms read timeout, fall back to empty listCode pathMerged, and a test with a stalled recs stub asserts 200 within the gateway budget
Cap concurrent recs calls per podFailure modeLoad test with recs at 8 s p99 keeps cart p99 under target
Audit client timeouts against gateway route timeoutsFailure modeNo client timeout on a user request path exceeds its route timeout
Staged flag rollout with automatic haltTriggerFlag service enforces stages for flags tagged request-path
Alert on cart 5xx and latency SLO burnDetectionAlert fires in a replay of this incident’s metrics within 3 minutes

Two more checks help. First, ask whether the action items would have prevented a different trigger through the same code path, such as a recommendations deploy instead of a flag. If yes, you have hardened the path rather than patched the event. Second, look for the same pattern elsewhere. The cart was one endpoint; other endpoints may make the same kind of blocking call to optional services. That search is often the most valuable action item in the document. For a related discussion of why these patterns pass tests, see A PR Passed CI but Broke Production.

How do you review an RCA before publishing it?

Review an RCA the way you would review code: against a checklist, by someone who was not the author. The questions below catch most gaps.

Chain complete
  • Trigger has a timestamp and a source record
  • Code path names entry point, function, and commit
  • Failure mode explains why the path broke then
  • Impact is measured, bounded, and sourced
Evidence
  • Every claim links to a trace, log, metric, or diff
  • Hypotheses are labeled and in open questions
  • Impact counts say whether they are requests, sessions, or users
Actions
  • Each action names the link it breaks
  • Each has an owner, a due date, and a done-when
  • At least one hardens the code path, not just the trigger
  • Someone searched for the same pattern elsewhere

If the incident involved an AI coding agent or an automated change, the same chain applies with one more link to document: the agent’s actions and tool calls that produced the change. The postmortem template for incidents an AI agent caused covers that extra section.

How Tomosu helps

The hardest section of an RCA to write is the code path, because it needs context that is spread across the repository: which endpoints reach a function, what timeouts and retries sit around a call, and which change introduced the behavior. Tomosu analyzes the codebase and each change for production reliability risk, which gives you that context before and after an incident:

These signals roll up into the Production Reliability Index. The goal is modest: fewer incidents where the code path section is the first time anyone traced the path.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

What is the difference between a trigger and a root cause in an RCA?

The trigger is the event that started the incident, such as a deploy, a feature flag change, a config push, or a traffic spike. The root cause is the latent weakness the trigger exposed, such as a blocking call with no fallback or a timeout longer than its caller’s. The same trigger on a well-built code path would cause little or no harm, which is why an RCA should name both.

What should an RCA include?

A one-line summary of the chain, a measured user impact statement, the causal chain (trigger, code path, failure mode, user impact), contributing factors with evidence, why tests, review, rollout, and alerting did not catch it, a timeline as supporting evidence, what went well and where you got lucky, action items tied to links in the chain, and open questions.

How do you describe user impact in an RCA?

Say who was affected, what they experienced, how many were affected, from when to when, and whether any data must be repaired, with the source for each number. Measure at the edge or with real user monitoring where possible, and state whether counts are requests, sessions, or users, because retries and reloads inflate request counts.

Is the 5 whys method enough for a software RCA?

It is a useful prompt, but on its own it tends to produce a single linear story that stops at whichever answer the group agrees on, often a person. Software incidents usually have several contributing conditions. Use the questions to explore, then write the result as a causal chain with evidence and a list of contributing factors.

How detailed should the code path section of an RCA be?

Detailed enough that an engineer who was not involved can follow it in the repository: the entry point users hit, each function or service hop, the call that failed or blocked, the commit or pull request that introduced the behavior, and the facts from outside the diff, such as worker counts or gateway timeouts, that made it fail.

What makes a good RCA action item?

A good action item names the link in the causal chain it breaks, has a single owner and a due date, and has a verifiable done-when condition, such as a test that fails before the fix and passes after. Items like “be more careful” or “improve monitoring” break no specific link and rarely change anything.

Who should write the RCA, and how soon?

The people closest to the incident, usually the incident commander and the engineers who owned the affected code, should draft it while evidence and memory are fresh, typically within a few working days. Someone who was not involved should review it. Keep it blameless: describe what systems and processes allowed, not who made a mistake.


The best RCAs read like a proof: this event, through this code, caused this harm, and these changes break the chain. Tomosu maps the code paths so that proof is easier to write, and so fewer incidents need one. Assess your repository →