Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Suspect commits

Which Commit Caused the Production Incident? A Practical Investigation Workflow

Tomosu AI·15 min read·

Errors spiked twenty minutes after a deploy, and everyone in the incident channel has the same question: which commit did this? The fastest way to find the commit that introduced a bug is not to read the whole diff. It is to shrink the set of suspects with evidence you already have, then confirm one of them with a test you can repeat.

Quick answer

To find the commit that caused a production incident, work from time to code. Pin the first moment things went bad, map it to the deploy (and any config or flag change) that preceded it, list the commits between the last good and first bad deploys, keep the ones that touch the failing code path, then confirm the top suspect by reverting it or with git bisect.

Every step removes candidates. A year of history becomes one deploy window, the window becomes the pull requests that touched the failing code, and those become a ranked list you can test. Along the way you also have to rule out the changes that never appear in git log: feature flags, configuration, dependencies, and infrastructure.

How do you find the commit that introduced a bug?

Treat it as a search problem with a shrinking candidate set. You start with every change that could matter and apply one filter at a time, each backed by a different piece of evidence. The filters are cheap at the top (a timestamp, a deploy record) and more expensive at the bottom (a reproduction, a revert). Doing them in that order is what keeps an incident investigation from turning into a line-by-line read of a 40-commit release.

FROM ALL COMMITS TO ONE CAUSE (ILLUSTRATIVE) All commits on main Thousands of changes, most irrelevant Commits in the deploy window 38 commits between two SHAs PRs touching the failing path 4 pull requests Ranked suspects 2 PRs, strongest first Confirmed cause 1 commit START Code, config, flags, deps, infra FILTER BY TIME First bad time, deploy SHAs FILTER BY CODE PATH Stack frames, changed files RANK Change type, blast radius CONFIRM Revert, flag off, git bisect
Each stage uses a different piece of evidence. The cheap filters come first, so the expensive confirmation step only runs on one or two candidates.

Two principles make this work. First, anchor on SHAs, not dates. A commit’s date says when it was written or rebased, not when it was merged or deployed. Second, keep a written list of every candidate and why it was kept or dropped. That list is what stops the team from chasing the most recent or most suspicious-looking commit instead of the one the evidence points to.

When did the incident actually start?

The alert time is when a threshold was crossed, often several minutes after the problem began. You want the first bad time: the earliest moment the failing signal departs from its baseline. Get it from the most specific signal you have:

Then overlay every change event on the same time axis: deploys, feature flag changes, config pushes, dependency or base image updates, infrastructure changes, and traffic shifts. The relationship between the first bad time and those events already tells you a lot.

FIRST BAD TIME VS CHANGE EVENTS SUSPECT WINDOW first bad 14:42 13:0014:0015:0016:00 13:10 DEPLOY r41 a1b2c3d Last good 14:05 DEPLOY r42 9f8e7d6 No change in errors 14:40 FLAG 50% checkout_v2 Not in git log 15:20 FLAG OFF Errors recover Mitigated
The deploy did not move the error rate. The flag change 35 minutes later did. The suspect is still code from r42, but only the code behind checkout_v2.

Read the gap between the change and the first bad time carefully:

Which deploy was running, and which SHA is it?

“We deployed around 2 pm” is not good enough. You need the exact git SHA of the last deploy known to be healthy and of the first deploy known to be bad. If your build embeds that SHA in its artifacts, the lookup takes a minute.

FROM WHAT IS RUNNING TO WHAT CHANGED WORKLOAD Pod, VM, or function version IMAGE Tag and digest sha256:4be1… REVISION OCI label or build info GIT SHA 9f8e7d6 First bad deploy PULL REQUESTS Merged since last good SHA $ git log --oneline --first-parent a1b2c3d..9f8e7d6 9f8e7d6 Merge pull request #2141 from feat/checkout-v2-totals 77c01ab Merge pull request #2138 from chore/bump-http-client 5d9e320 Merge pull request #2135 from fix/tax-rounding … one line per merged PR
If every artifact records the commit it was built from, “which code was running?” becomes a lookup instead of an argument.

Common places to find the SHA:

Shell · what is running?
# Kubernetes: image currently in the Deployment, and the rollout history
kubectl get deployment orders -o jsonpath='{.spec.template.spec.containers[0].image}'
kubectl rollout history deployment/orders

# Image label written at build time (OCI standard annotation)
docker inspect --format '{{ index .Config.Labels "org.opencontainers.image.revision" }}' \
  registry.example.com/orders:r42

# Build info exposed by the app itself, e.g. Spring Boot Actuator with git info
curl -s https://orders.internal/actuator/info | jq '.git.commit.id'

# Git tags, if CI tags each release
git rev-parse --short r41^{commit} r42^{commit}

If none of those exist, your CI system’s deploy job history is the fallback. Fix the gap after the incident: stamp the SHA into the image label, the binary (Go records vcs.revision in its build info since Go 1.18), a /version endpoint, and your error tracker’s release field. Many error trackers and APM tools also accept release or deploy markers, so the first bad time and the version appear on the same chart.

Check what was actually deployed

A tag can be moved, a rollout can stall halfway, and a canary can run a different version from the rest of the fleet. Confirm the SHA from the running workload or from telemetry labelled with the version, not from the release notes.

List every commit between last good and first bad

With two SHAs you have a closed range. git log A..B lists commits reachable from B but not from A: exactly what the bad deploy contains that the good one did not.

Shell · the deploy window
GOOD=a1b2c3d   # last good deploy
BAD=9f8e7d6    # first bad deploy

# One line per merged PR (merge commits or squash commits on main)
git log --oneline --first-parent $GOOD..$BAD

# Every commit, including those inside merged branches
git log --oneline $GOOD..$BAD

# Which files changed, and how much
git diff --stat $GOOD $BAD

# Find the PR for a commit (GitHub CLI, or the REST endpoint)
gh pr list --state merged --search "5d9e320"
gh api repos/{owner}/{repo}/commits/5d9e320/pulls --jq '.[].html_url'

--first-parent follows only the main line, so with merge commits or squash merges each line is one PR. That is the right unit for an incident: PRs have descriptions, reviewers, and linked tickets, and a revert usually targets a whole PR.

If you only have a time window and no SHAs, git log --since and --until filter on commit dates, which are set when a commit is created or rebased. A PR opened last week and merged an hour before the deploy can fall outside a date filter. Resolve the deploys to SHAs first.

How do you match a production error to a suspect commit?

Now cut the range down to the PRs that touch the failing code path. The stack trace, the failing endpoint, and the trace spans tell you which files and functions were involved. Intersect those with what changed.

Shell · narrow by code path
# Commits in the window that touched the files in the stack trace
git log --oneline $GOOD..$BAD -- src/checkout/CartTotals.java src/pricing/

# Pickaxe: commits that added or removed a string (a field, a flag name)
git log --oneline -S'discountCents' $GOOD..$BAD

# Commits whose diff lines match a regex
git log --oneline -G'checkout_v2' $GOOD..$BAD

# History of one function (needs a hunk header git can recognise)
git log -L :applyDiscount:src/pricing/Discounts.java $GOOD..$BAD

# Who last changed the lines in the failing frame, ignoring bulk reformat commits
git blame -L 118,140 --ignore-revs-file .git-blame-ignore-revs $BAD -- src/checkout/CartTotals.java

Three cautions keep this step honest:

  1. The frame that throws is not always the frame that changed. A NullPointerException in CartTotals can be caused by a serializer change that stopped populating a field two calls earlier. Look at every application frame in the trace, and at the callers of the changed code.
  2. git blame shows the last change to a line, not the change that broke it. The bug may be a line that was deleted, or a behaviour change in a function the line calls.
  3. Shared code widens the net. A change to a serializer, a retry wrapper, or a shared client touches the failing path without appearing in its stack. This is where knowing the blast radius of each change earns its keep.

Error trackers automate part of this. Sentry, for example, shows suspect commits on an issue once releases and a repository integration are set up, by relating the files and lines in the stack trace to recent commits. Treat that as a strong lead, not a verdict: it can only see what the stack shows, and it cannot see flags, configuration, or data.

Ranking what is left

Order the remaining PRs by how many independent signals point at each: overlap with the failing frames, timing that fits the first bad time, the kind of change (a query, a retry policy, a serialization format, a timeout, a dependency bump), and how many callers depend on the changed code. A PR that only renamed a test helper can be dropped; a PR that changed the default value of a timeout in a shared client stays near the top even if it is not in the stack.

What about changes that are not commits?

Plenty of incidents start with a change that git log will never show. If the first bad time does not line up with a deploy, or the suspect commits do not explain the symptom, check these before bisecting anything.

Change typeWhere to find itWhat it looks like
Feature flagFlag service audit logErrors step up at a ramp change, with no deploy; affected users match the targeting rule
Runtime configurationConfig repo, parameter store, ConfigMap historyBehaviour change on restart or reload; pods started before the change behave differently
Dependency versionLockfile diff between the two SHAsA commit exists, but its diff is one lockfile line hiding a large behaviour change
Base image or runtimeImage digest and Dockerfile historySame application SHA, different OS libraries, TLS roots, or language runtime
InfrastructureTerraform plan history, cloud audit logsLimits, instance types, network policy, or DNS changed around the first bad time
Schema or dataMigration logs, backfill jobsOnly some records fail; a migration or backfill ran near the first bad time
Upstream serviceTheir deploy markers and status pageYour code is unchanged; errors are in calls to one dependency
Shell · dependencies in the window
# Lockfile changes are commits too, but easy to skim past
git diff $GOOD $BAD -- package-lock.json go.sum poetry.lock gradle.lockfile

# Commits in the window that touched dependency manifests, at any depth
git log --oneline $GOOD..$BAD -- ':(glob)**/package.json' ':(glob)**/go.mod' ':(glob)**/pom.xml'

How do you use git bisect to find a production bug?

Once the range is small and you can reproduce the failure outside production, git bisect finds the first bad commit by binary search. It checks out the midpoint, you tell it good or bad, and it halves the range. Sixteen candidates take four tests; a thousand take about ten.

BINARY SEARCH OVER 16 COMMITS 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 COMMITS c1 good, c16 bad STEP 1 c8 is good STEP 2 c12 is bad STEP 3 c10 is good STEP 4 c11 is bad c11 is the first bad commit. Four tests instead of up to fifteen. Shaded bars show the remaining range.
Bisect needs about log₂(N) tests. The expensive part is not the search; it is having a test that says good or bad reliably at every commit.

Start with the SHAs you already have: the first bad deploy and the last good one. Then automate the verdict with git bisect run:

Shell · automated bisect
git bisect start 9f8e7d6 a1b2c3d          # bad first, then good
git bisect run /tmp/repro-checkout.sh     # script lives outside the repo

# ...bisect prints: <sha> is the first bad commit
git bisect log > bisect-incident-2141.log  # keep it for the postmortem
git bisect reset                          # return to where you started
/tmp/repro-checkout.shexit codes matter
#!/usr/bin/env bash
# 0 = good, 1-127 (except 125) = bad, 125 = cannot test this commit
./gradlew -q compileJava || exit 125     # broken build: skip, don't blame

# Replay the production input that fails, captured from the incident
./gradlew -q test --tests 'com.example.checkout.CartTotalsReplayTest' \
  -Dreplay.file=/tmp/incident-2141-cart.json

The details that make or break an automated bisect:

When git bisect does not work

Bisect assumes one reproducible, deterministic failure that appears at one commit and stays. Production incidents often break that assumption.

CAN BISECT ANSWER THIS ONE? Q1 · REPRODUCTION Can you trigger it outside production? Q2 · DETERMINISM Does it fail every time it runs? Q3 · TESTABLE HISTORY Does every commit in the range build? Narrow by evidence Flag off or canary revert Repeat each step Bad if any run of N fails Skip untestable commits exit 125, or --first-parent NONONO YESYES git bisect run with the reproduction script: fast, mechanical, and repeatable.
Bisect is the confirmation tool of choice when a reliable reproduction exists. Without one, confirmation moves into production with controlled, reversible changes.

Should you revert the suspect commit or roll forward?

You do not need certainty to mitigate. You need a reversible action with an acceptable blast radius. Often the revert is also the confirmation: if errors drop back to the baseline when only one change is undone, you have strong evidence for that change.

OptionChoose it whenWatch out for
Roll back the deployThe whole release is suspect and the previous artifact is known goodSchema migrations that ran forward; other fixes in the release that users now depend on
Turn off a flagThe first bad time matches a flag change and the old path is still in the codeFlags that also gate data writes; the old path may not handle new data
Revert one PREvidence points to one PR and the rest of the release is healthyRevert conflicts; for merge commits use git revert -m 1 <merge-sha>, and re-merging later needs the revert reverted
Roll forward with a fixRollback is unsafe or impossible, and the fix is small and well understoodA rushed fix is a new untested change under pressure; keep it minimal
Mitigate first, attribute second

Rolling back a release restores service without telling you which commit caused the problem. That is fine. Keep the bad SHA, the evidence, and the reproduction, and finish the attribution after users are no longer affected. The postmortem needs the answer; the incident only needs the bleeding to stop.

Build the evidence table

Write the suspects down with the evidence for and against each. It turns an argument in the incident channel into a list anyone can check, and it is the core of the postmortem’s “trigger” section.

CandidateIn windowTouches failing pathTimingOther signalsVerdict
PR #2141 checkout v2 totalsYes (r42)Yes: CartTotals in every traceDormant until flag 14:40Errors only for flag cohortPrimary suspect; confirmed by flag off
Flag checkout_v2 50%Not a commitActivates #21412 min before first badRecovery at flag offTrigger
PR #2138 HTTP client bumpYes (r42)Indirectly: shared clientNo step at 14:05No timeouts in tracesDropped
PR #2135 tax roundingYes (r42)Same package, not in traceNo step at 14:05Unit tests cover changeDropped
DB index migrationRan 13:55No47 min earlyQuery latency flatDropped

An illustrative evidence table. Record the reason for every dropped candidate, not only the winner.

Note how the answer has two parts: the commit that contained the defect and the change that exposed it. Both belong in the postmortem, and both are useful for preventing the next one. This is also why automated root cause analysis is less about time to resolve and more about whether the evidence chain from change to symptom was recorded.

How Tomosu helps

The slow part of this workflow is the middle of the funnel: going from 38 commits in a deploy window to the few PRs that actually touch the failing code path, and knowing which of those can affect it indirectly through shared code. Tomosu scans repositories and pull requests and keeps that mapping ready before the incident starts:

The same signals feed the Production Reliability Index. Code Volatility, Deployment Velocity, and Fragility Index describe how much a codebase changes and how exposed those changes are, which is the context an investigator needs when a release contains dozens of PRs. And the PRs that would sit at the top of the suspect list are the ones worth flagging before merge.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

How do I find which commit caused a production incident?

Establish the first bad time from error rates and metrics, map it to the deploy that preceded it, and resolve the last good and first bad deploys to git SHAs. List the commits between them with git log, keep the ones that touch the failing stack frames or code path, rank them by evidence, and confirm the top suspect with a revert, a flag change, or git bisect.

How does git bisect find the commit that introduced a bug?

git bisect performs a binary search between a known bad commit and a known good commit. At each step it checks out the midpoint, you or a script mark it good or bad, and the range halves. With git bisect run, the script’s exit code decides: 0 is good, 1 to 127 except 125 is bad, and 125 skips a commit that cannot be tested.

What is a suspect commit?

A suspect commit is a change in the deploy window that plausibly caused an error, usually because it modified files or lines that appear in the failing stack trace. Some error tracking tools, such as Sentry, can show suspect commits on an issue when releases and a repository integration are configured. A suspect is a lead to confirm, not a proven cause.

Can I use git bisect if the bug only happens in production?

Only if you can build a test that reproduces it, for example by replaying a captured request or record against each commit. If the failure depends on production data, traffic, or concurrency you cannot reproduce, narrow the suspects with evidence and confirm in production with a controlled change, such as turning off a flag or reverting one PR on a canary.

How do I handle commits that do not build during git bisect?

Mark them as untestable with git bisect skip, or have your git bisect run script exit with code 125 when the build fails. Bisect then chooses a nearby commit instead. Using git bisect start --first-parent, available since Git 2.29, also avoids work-in-progress commits inside merged branches.

What if the incident was caused by a feature flag or config change, not a commit?

Then git log will not show the trigger. Compare the first bad time with flag audit logs, configuration history, dependency and base image changes, and infrastructure changes. Often the answer has two parts: the commit that contained the defect and the flag or config change that made it reachable.

Should I revert the suspect commit or roll forward with a fix?

Prefer the smallest reversible action that restores service: turn off a flag, revert one PR, or roll back the release. Roll forward only when rollback is unsafe, for example after an irreversible migration, and the fix is small and well understood. A clean recovery after reverting one change is also strong evidence that the change was the cause.


The commit that caused an incident was usually a pull request someone reviewed. Tomosu maps which code paths each change touches, so the suspect list is short when it matters. Assess your repository →