Errors spiked twenty minutes after a deploy, and everyone in the incident channel has the same question: which commit did this? The fastest way to find the commit that introduced a bug is not to read the whole diff. It is to shrink the set of suspects with evidence you already have, then confirm one of them with a test you can repeat.
To find the commit that caused a production incident, work from time to code. Pin the first moment things went bad, map it to the deploy (and any config or flag change) that preceded it, list the commits between the last good and first bad deploys, keep the ones that touch the failing code path, then confirm the top suspect by reverting it or with git bisect.
- Time: first bad timestamp from error rates and metrics, not from the alert.
- Deploy: the running image or build maps to a git SHA.
- Range:
git log last_good..first_badis the full list of code suspects. - Path: stack frames and changed files narrow it to a few PRs.
- Confirm: revert, flag off, or
git bisect runwith a reproducible test.
Every step removes candidates. A year of history becomes one deploy window, the window becomes the pull requests that touched the failing code, and those become a ranked list you can test. Along the way you also have to rule out the changes that never appear in git log: feature flags, configuration, dependencies, and infrastructure.
How do you find the commit that introduced a bug?
Treat it as a search problem with a shrinking candidate set. You start with every change that could matter and apply one filter at a time, each backed by a different piece of evidence. The filters are cheap at the top (a timestamp, a deploy record) and more expensive at the bottom (a reproduction, a revert). Doing them in that order is what keeps an incident investigation from turning into a line-by-line read of a 40-commit release.
Two principles make this work. First, anchor on SHAs, not dates. A commit’s date says when it was written or rebased, not when it was merged or deployed. Second, keep a written list of every candidate and why it was kept or dropped. That list is what stops the team from chasing the most recent or most suspicious-looking commit instead of the one the evidence points to.
When did the incident actually start?
The alert time is when a threshold was crossed, often several minutes after the problem began. You want the first bad time: the earliest moment the failing signal departs from its baseline. Get it from the most specific signal you have:
- Error tracking: the first-seen timestamp of the new error group, plus which release it was first seen in.
- Metrics: error rate, latency percentiles, or saturation for the affected endpoint, zoomed in until you can see the step.
- Logs: the first occurrence of the new exception or log message, ideally split by host or pod.
- Traces: the first failing span for the affected route, which also tells you which service and version served it.
Then overlay every change event on the same time axis: deploys, feature flag changes, config pushes, dependency or base image updates, infrastructure changes, and traffic shifts. The relationship between the first bad time and those events already tells you a lot.
checkout_v2.Read the gap between the change and the first bad time carefully:
- Errors start within seconds or minutes of a rollout and grow as pods are replaced: the new code is a strong suspect.
- Errors start later, at a flag flip or config push: the code may have shipped earlier and been dormant.
- Errors start hours later with no change event: look for load, data shape, a scheduled job, cache expiry, a certificate, or a dependency outage. A slow resource leak from a deploy can also take hours to show; memory growth before an OOMKilled pod is the classic example.
- Errors appear only on new pods during a rolling deploy: split the metric by version label. If old pods are clean and new pods fail, you have confirmed the release before reading any code.
Which deploy was running, and which SHA is it?
“We deployed around 2 pm” is not good enough. You need the exact git SHA of the last deploy known to be healthy and of the first deploy known to be bad. If your build embeds that SHA in its artifacts, the lookup takes a minute.
Common places to find the SHA:
# Kubernetes: image currently in the Deployment, and the rollout history
kubectl get deployment orders -o jsonpath='{.spec.template.spec.containers[0].image}'
kubectl rollout history deployment/orders
# Image label written at build time (OCI standard annotation)
docker inspect --format '{{ index .Config.Labels "org.opencontainers.image.revision" }}' \
registry.example.com/orders:r42
# Build info exposed by the app itself, e.g. Spring Boot Actuator with git info
curl -s https://orders.internal/actuator/info | jq '.git.commit.id'
# Git tags, if CI tags each release
git rev-parse --short r41^{commit} r42^{commit}
If none of those exist, your CI system’s deploy job history is the fallback. Fix the gap after the incident: stamp the SHA into the image label, the binary (Go records vcs.revision in its build info since Go 1.18), a /version endpoint, and your error tracker’s release field. Many error trackers and APM tools also accept release or deploy markers, so the first bad time and the version appear on the same chart.
A tag can be moved, a rollout can stall halfway, and a canary can run a different version from the rest of the fleet. Confirm the SHA from the running workload or from telemetry labelled with the version, not from the release notes.
List every commit between last good and first bad
With two SHAs you have a closed range. git log A..B lists commits reachable from B but not from A: exactly what the bad deploy contains that the good one did not.
GOOD=a1b2c3d # last good deploy
BAD=9f8e7d6 # first bad deploy
# One line per merged PR (merge commits or squash commits on main)
git log --oneline --first-parent $GOOD..$BAD
# Every commit, including those inside merged branches
git log --oneline $GOOD..$BAD
# Which files changed, and how much
git diff --stat $GOOD $BAD
# Find the PR for a commit (GitHub CLI, or the REST endpoint)
gh pr list --state merged --search "5d9e320"
gh api repos/{owner}/{repo}/commits/5d9e320/pulls --jq '.[].html_url'
--first-parent follows only the main line, so with merge commits or squash merges each line is one PR. That is the right unit for an incident: PRs have descriptions, reviewers, and linked tickets, and a revert usually targets a whole PR.
If you only have a time window and no SHAs, git log --since and --until filter on commit dates, which are set when a commit is created or rebased. A PR opened last week and merged an hour before the deploy can fall outside a date filter. Resolve the deploys to SHAs first.
How do you match a production error to a suspect commit?
Now cut the range down to the PRs that touch the failing code path. The stack trace, the failing endpoint, and the trace spans tell you which files and functions were involved. Intersect those with what changed.
# Commits in the window that touched the files in the stack trace
git log --oneline $GOOD..$BAD -- src/checkout/CartTotals.java src/pricing/
# Pickaxe: commits that added or removed a string (a field, a flag name)
git log --oneline -S'discountCents' $GOOD..$BAD
# Commits whose diff lines match a regex
git log --oneline -G'checkout_v2' $GOOD..$BAD
# History of one function (needs a hunk header git can recognise)
git log -L :applyDiscount:src/pricing/Discounts.java $GOOD..$BAD
# Who last changed the lines in the failing frame, ignoring bulk reformat commits
git blame -L 118,140 --ignore-revs-file .git-blame-ignore-revs $BAD -- src/checkout/CartTotals.java
Three cautions keep this step honest:
- The frame that throws is not always the frame that changed. A
NullPointerExceptioninCartTotalscan be caused by a serializer change that stopped populating a field two calls earlier. Look at every application frame in the trace, and at the callers of the changed code. git blameshows the last change to a line, not the change that broke it. The bug may be a line that was deleted, or a behaviour change in a function the line calls.- Shared code widens the net. A change to a serializer, a retry wrapper, or a shared client touches the failing path without appearing in its stack. This is where knowing the blast radius of each change earns its keep.
Error trackers automate part of this. Sentry, for example, shows suspect commits on an issue once releases and a repository integration are set up, by relating the files and lines in the stack trace to recent commits. Treat that as a strong lead, not a verdict: it can only see what the stack shows, and it cannot see flags, configuration, or data.
Ranking what is left
Order the remaining PRs by how many independent signals point at each: overlap with the failing frames, timing that fits the first bad time, the kind of change (a query, a retry policy, a serialization format, a timeout, a dependency bump), and how many callers depend on the changed code. A PR that only renamed a test helper can be dropped; a PR that changed the default value of a timeout in a shared client stays near the top even if it is not in the stack.
What about changes that are not commits?
Plenty of incidents start with a change that git log will never show. If the first bad time does not line up with a deploy, or the suspect commits do not explain the symptom, check these before bisecting anything.
| Change type | Where to find it | What it looks like |
|---|---|---|
| Feature flag | Flag service audit log | Errors step up at a ramp change, with no deploy; affected users match the targeting rule |
| Runtime configuration | Config repo, parameter store, ConfigMap history | Behaviour change on restart or reload; pods started before the change behave differently |
| Dependency version | Lockfile diff between the two SHAs | A commit exists, but its diff is one lockfile line hiding a large behaviour change |
| Base image or runtime | Image digest and Dockerfile history | Same application SHA, different OS libraries, TLS roots, or language runtime |
| Infrastructure | Terraform plan history, cloud audit logs | Limits, instance types, network policy, or DNS changed around the first bad time |
| Schema or data | Migration logs, backfill jobs | Only some records fail; a migration or backfill ran near the first bad time |
| Upstream service | Their deploy markers and status page | Your code is unchanged; errors are in calls to one dependency |
# Lockfile changes are commits too, but easy to skim past
git diff $GOOD $BAD -- package-lock.json go.sum poetry.lock gradle.lockfile
# Commits in the window that touched dependency manifests, at any depth
git log --oneline $GOOD..$BAD -- ':(glob)**/package.json' ':(glob)**/go.mod' ':(glob)**/pom.xml'
How do you use git bisect to find a production bug?
Once the range is small and you can reproduce the failure outside production, git bisect finds the first bad commit by binary search. It checks out the midpoint, you tell it good or bad, and it halves the range. Sixteen candidates take four tests; a thousand take about ten.
Start with the SHAs you already have: the first bad deploy and the last good one. Then automate the verdict with git bisect run:
git bisect start 9f8e7d6 a1b2c3d # bad first, then good
git bisect run /tmp/repro-checkout.sh # script lives outside the repo
# ...bisect prints: <sha> is the first bad commit
git bisect log > bisect-incident-2141.log # keep it for the postmortem
git bisect reset # return to where you started
#!/usr/bin/env bash
# 0 = good, 1-127 (except 125) = bad, 125 = cannot test this commit
./gradlew -q compileJava || exit 125 # broken build: skip, don't blame
# Replay the production input that fails, captured from the incident
./gradlew -q test --tests 'com.example.checkout.CartTotalsReplayTest' \
-Dreplay.file=/tmp/incident-2141-cart.json
The details that make or break an automated bisect:
- Exit codes.
0marks the commit good,1to127mark it bad,125skips it, and anything above127aborts the bisect. A build failure must exit125, or bisect will blame the wrong commit. - Keep the test outside the tree. Bisect checks out old commits. A new test file committed today does not exist at those commits, so keep the script and fixture in a path like
/tmp, or as untracked files git will not touch. - Find the PR first.
git bisect start --first-parent(Git 2.29 and later) walks only the main line, which gives you the bad PR. Bisect inside that PR afterwards if you need the exact commit. - Custom terms. When you are hunting a performance change,
--term-old=fast --term-new=slowreads better than good and bad.
When git bisect does not work
Bisect assumes one reproducible, deterministic failure that appears at one commit and stays. Production incidents often break that assumption.
- Production-only data. The bug needs a customer record, a traffic mix, or a data volume you cannot copy. Try capturing the failing request or record (scrubbed of personal data) as a replay fixture first. If that fails, confirm with a controlled change in production instead: flag off, or revert the top suspect on a canary and compare its error rate with the rest of the fleet.
- Flaky or load-dependent failures. A race or a timeout under concurrency may pass by luck. Run the test several times per step and exit non-zero if any run fails, or bisect on a measured error rate under a fixed load. A false “good” sends the search into the wrong half and the answer will be wrong without any warning.
- Interaction bugs. Two commits are each fine alone and fail together. Bisect reports the second one, which is correct but incomplete. The evidence table below should list both.
- The cause is not a commit. If the same SHA is healthy before a flag, config, or infrastructure change and broken after it, bisecting code will only find the commit that made the code reachable, not the change that triggered it.
Should you revert the suspect commit or roll forward?
You do not need certainty to mitigate. You need a reversible action with an acceptable blast radius. Often the revert is also the confirmation: if errors drop back to the baseline when only one change is undone, you have strong evidence for that change.
| Option | Choose it when | Watch out for |
|---|---|---|
| Roll back the deploy | The whole release is suspect and the previous artifact is known good | Schema migrations that ran forward; other fixes in the release that users now depend on |
| Turn off a flag | The first bad time matches a flag change and the old path is still in the code | Flags that also gate data writes; the old path may not handle new data |
| Revert one PR | Evidence points to one PR and the rest of the release is healthy | Revert conflicts; for merge commits use git revert -m 1 <merge-sha>, and re-merging later needs the revert reverted |
| Roll forward with a fix | Rollback is unsafe or impossible, and the fix is small and well understood | A rushed fix is a new untested change under pressure; keep it minimal |
Rolling back a release restores service without telling you which commit caused the problem. That is fine. Keep the bad SHA, the evidence, and the reproduction, and finish the attribution after users are no longer affected. The postmortem needs the answer; the incident only needs the bleeding to stop.
Build the evidence table
Write the suspects down with the evidence for and against each. It turns an argument in the incident channel into a list anyone can check, and it is the core of the postmortem’s “trigger” section.
| Candidate | In window | Touches failing path | Timing | Other signals | Verdict |
|---|---|---|---|---|---|
| PR #2141 checkout v2 totals | Yes (r42) | Yes: CartTotals in every trace | Dormant until flag 14:40 | Errors only for flag cohort | Primary suspect; confirmed by flag off |
Flag checkout_v2 50% | Not a commit | Activates #2141 | 2 min before first bad | Recovery at flag off | Trigger |
| PR #2138 HTTP client bump | Yes (r42) | Indirectly: shared client | No step at 14:05 | No timeouts in traces | Dropped |
| PR #2135 tax rounding | Yes (r42) | Same package, not in trace | No step at 14:05 | Unit tests cover change | Dropped |
| DB index migration | Ran 13:55 | No | 47 min early | Query latency flat | Dropped |
An illustrative evidence table. Record the reason for every dropped candidate, not only the winner.
Note how the answer has two parts: the commit that contained the defect and the change that exposed it. Both belong in the postmortem, and both are useful for preventing the next one. This is also why automated root cause analysis is less about time to resolve and more about whether the evidence chain from change to symptom was recorded.
How Tomosu helps
The slow part of this workflow is the middle of the funnel: going from 38 commits in a deploy window to the few PRs that actually touch the failing code path, and knowing which of those can affect it indirectly through shared code. Tomosu scans repositories and pull requests and keeps that mapping ready before the incident starts:
- Code paths per change: for each PR, the endpoints, jobs, and shared components its changed files sit on, so matching a stack trace to the deploy window is a lookup rather than a guess.
- Changes that reach the failing path indirectly: a PR that modifies a shared client, serializer, or retry policy is surfaced even when it is not in the stack trace.
- Ranking signals: findings weighed by blast radius and change type, so a timeout change on the checkout path ranks above a test refactor in the same release.
- The evidence behind each ranking: which files, which paths, and which risk findings the PR carried at review time, ready to drop into the evidence table.
The same signals feed the Production Reliability Index. Code Volatility, Deployment Velocity, and Fragility Index describe how much a codebase changes and how exposed those changes are, which is the context an investigator needs when a release contains dozens of PRs. And the PRs that would sit at the top of the suspect list are the ones worth flagging before merge.
Scan your repository with Tomosu →
Key takeaways
- Find the commit that introduced a bug by shrinking the candidate set: time, then deploy, then code path, then confirmation.
- Use the first bad time from metrics and errors, not the alert time, and overlay every change event on it.
- Resolve deploys to exact SHAs.
git log --first-parent good..badgives one line per PR in the window. - Match stack frames and changed files with path filters,
-S,-G,-L, andgit blame, but remember shared code that is not in the trace. - Rule out flags, config, dependencies, infrastructure, and data before bisecting code.
git bisect runis fast and reliable with a deterministic reproduction; build failures must exit125.- Mitigate with the smallest reversible action, and record every suspect with the reason it was kept or dropped.
Frequently asked questions
How do I find which commit caused a production incident?
Establish the first bad time from error rates and metrics, map it to the deploy that preceded it, and resolve the last good and first bad deploys to git SHAs. List the commits between them with git log, keep the ones that touch the failing stack frames or code path, rank them by evidence, and confirm the top suspect with a revert, a flag change, or git bisect.
How does git bisect find the commit that introduced a bug?
git bisect performs a binary search between a known bad commit and a known good commit. At each step it checks out the midpoint, you or a script mark it good or bad, and the range halves. With git bisect run, the script’s exit code decides: 0 is good, 1 to 127 except 125 is bad, and 125 skips a commit that cannot be tested.
What is a suspect commit?
A suspect commit is a change in the deploy window that plausibly caused an error, usually because it modified files or lines that appear in the failing stack trace. Some error tracking tools, such as Sentry, can show suspect commits on an issue when releases and a repository integration are configured. A suspect is a lead to confirm, not a proven cause.
Can I use git bisect if the bug only happens in production?
Only if you can build a test that reproduces it, for example by replaying a captured request or record against each commit. If the failure depends on production data, traffic, or concurrency you cannot reproduce, narrow the suspects with evidence and confirm in production with a controlled change, such as turning off a flag or reverting one PR on a canary.
How do I handle commits that do not build during git bisect?
Mark them as untestable with git bisect skip, or have your git bisect run script exit with code 125 when the build fails. Bisect then chooses a nearby commit instead. Using git bisect start --first-parent, available since Git 2.29, also avoids work-in-progress commits inside merged branches.
What if the incident was caused by a feature flag or config change, not a commit?
Then git log will not show the trigger. Compare the first bad time with flag audit logs, configuration history, dependency and base image changes, and infrastructure changes. Often the answer has two parts: the commit that contained the defect and the flag or config change that made it reachable.
Should I revert the suspect commit or roll forward with a fix?
Prefer the smallest reversible action that restores service: turn off a flag, revert one PR, or roll back the release. Roll forward only when rollback is unsafe, for example after an irreversible migration, and the fix is small and well understood. A clean recovery after reverting one change is also strong evidence that the change was the cause.
The commit that caused an incident was usually a pull request someone reviewed. Tomosu maps which code paths each change touches, so the suspect list is short when it matters. Assess your repository →