A merge gate that blocks on every reliability finding gets bypassed within a month. A gate that blocks on nothing gets ignored within a week. The useful policy sits between those two, and it needs a way to decide, finding by finding, which ones are worth stopping a merge for.
A reliability finding should block a merge only when the pull request introduced or activated it, the finding is verified with evidence, and its severity is high: likely data loss, an outage, or serious degradation on a critical path. Everything else should warn or be tracked.
- S1, critical: block. Emergency override only.
- S2, high: block unless a named approver records a compensating control.
- S3, medium: warn; the author acknowledges or files a ticket.
- S4, low: track, never block.
- Pre-existing or unverified: never block on its own.
This guide gives you a severity framework you can adopt as written or adjust: the questions to ask before severity, a rubric for rating impact and likelihood, the evidence a blocking finding must carry, an override path that does not become a loophole, and a rollout plan that starts in report-only mode. It applies whether the finding comes from a human reviewer, a static analyzer, or an AI review tool.
Why do all-or-nothing merge gates fail?
All-or-nothing merge gates fail because reliability findings vary by orders of magnitude in consequence, and a gate that treats them the same forces people to work around it. A merge gate is a condition that must be satisfied before a pull request can merge, usually a required status check enforced through branch protection or a ruleset (see GitHub’s documentation on protected branches).
- Block on everything and the gate fails on a missing jitter in a nightly job as loudly as on a non-idempotent payment retry. Engineers learn that red usually means noise. Admins get asked for bypasses, the check gets made optional “temporarily”, and the one finding that mattered merges with the rest.
- Block on nothing and findings become comments. Comments on a busy PR get resolved without being read. The finding is technically visible and practically absent.
- Block on a score threshold without saying what the score means, and you get the problem described in A Signal Is Not a Gate: a measurement promoted to a decision it was never built to make.
The fix is not a better threshold. It is a policy that separates three questions people usually collapse into one: is this finding the PR’s, is it real, and how bad would it be?
What should you ask before rating severity?
Before rating severity, ask whether the pull request owns the finding and whether the finding is verified. Only findings that pass both questions should reach the severity rubric. This ordering keeps old debt and speculative findings from ever blocking a merge, which is most of what makes gates feel unfair.
Question 1: Does the pull request own the finding?
A finding belongs to a pull request if the PR introduced it, or if the PR activated it: the risk was already in the code, but the PR makes it reachable, more frequent, or more severe. A missing timeout in an old client that the PR starts calling from checkout is activated. The same missing timeout in a client nothing new touches is pre-existing, and it goes to a backlog. The distinction, and why it needs a view beyond the diff, is covered in What a Repository-Wide Review Finds That a PR Diff Misses.
Question 2: Is the finding verified?
A blocking finding must be checkable by someone other than its author. At minimum it should name the file and line, the condition that triggers the failure, and the impact. “This could cause issues under load” is not verifiable. “place_order holds a database connection and row locks while get_price makes an HTTP call with up to three attempts; ten slow checkouts exhaust the pool” is. When a finding lacks this, the right action is a warning that asks for the evidence, not a block. For how to check a finding quickly, see AI Code Review False Positives: How to Verify a Reliability Finding.
How do you rate the severity of a reliability finding?
Rate the severity of a reliability finding by combining its impact if it happens with the likelihood that production conditions trigger it, then raising it one level if it sits on a critical or shared path, or if its effect cannot be undone. The result maps to one of four levels, each with a fixed action.
| Dimension | Levels | How to decide |
|---|---|---|
| Impact | Minor, Major, Severe | Severe: data loss or corruption, duplicate charges, full outage, security exposure. Major: errors or high latency on a critical path, recoverable. Minor: a non-critical feature degrades. |
| Likelihood | Rare, Plausible, Likely | Likely: normal traffic triggers it. Plausible: needs a traffic peak, a slow dependency, or a retry. Rare: needs an unusual combination, such as a narrow race plus a failover. |
| Blast radius | +1 level if wide | The code is on a critical user path (checkout, login, payments) or uses a shared resource (connection pool, queue, cache) that many paths depend on. |
| Reversibility | +1 level if irreversible | The damage persists after rollback: data written, messages sent, emails delivered, a column dropped. |
The adjustments matter more than they look. A retry added to an internal reporting job and the same retry added to the payment client have the same impact and likelihood in isolation. What separates them is the path they sit on and whether a retried call can charge a card twice. For how to size that reach, see How to Assess the Blast Radius of a Code Change.
Here is how common reliability findings land when the rubric is applied. Your own list should come from your incident history, because the fastest way to get agreement on a rubric is to rate past incidents with it.
| Finding | Impact | Likelihood | Adjust | Severity and action |
|---|---|---|---|---|
| Retry on a non-idempotent payment call with no idempotency key | Severe | Plausible | Irreversible | S1: block |
| Migration drops or rewrites a column still read by the running version | Severe | Likely | Irreversible | S1: block |
| HTTP call inside a DB transaction on the checkout path | Major | Plausible | Critical path | S1: block |
| Outbound call with no timeout on a request path | Major | Plausible | None | S2: block unless overridden |
| Unbounded query on a table that grows daily, admin page | Minor | Likely | None | S3: warn |
| Retry without jitter in a nightly batch job | Minor | Plausible | None | S4: track |
| Swallowed exception in a metrics exporter | Minor | Rare | None | S4: track |
Illustrative ratings. The same finding can rate differently in your system; that is the point of rating it.
A gate earns the right to block by being right almost every time it does.
Which reliability findings should never block a merge?
Findings that should never block a merge on their own are pre-existing findings the PR does not change, unverified findings, low-severity findings, and findings about style rather than behavior. Keeping these out of the blocking path is what lets the blocking path stay trusted.
| Finding type | Why it should not block | What to do instead |
|---|---|---|
| Pre-existing, exposure unchanged | Blocks an unrelated change for old debt; punishes whoever touched the file | Backlog item with an owner and a severity |
| Unverified or speculative | “May cause issues” cannot be checked or fixed with confidence | Warn and ask for file, line, and trigger |
| S4 or low confidence S3 | The cost of stopping the merge is higher than the risk | Comment or ticket |
| Style or naming | Not a reliability finding | Linter or ordinary review |
| Test coverage percentage alone | Coverage says what ran, not what is safe | Ask whether the changed failure paths are tested |
The most common way severity policies fail is severity inflation: tools and reviewers rate everything high because nobody is blamed for a false alarm. Require every S1 and S2 to carry the evidence from Question 2, and sample blocked findings each month to check they were real.
How do merge gate overrides work without becoming a loophole?
Merge gate overrides work when an override is a recorded decision by a named person, tied to a compensating control and a follow-up, rather than a button that makes the check go green. A compensating control is a measure that limits the damage if the finding turns out to be real: a default-off feature flag, a canary with automatic rollback, a rate limit, or a tested rollback plan.
- S1 overrides are for emergencies, such as a hotfix during an incident, and should need the same approval as the incident itself. If S1 overrides happen more than rarely, the rubric is rating too many things S1.
- S2 overrides are normal and expected. They say “we know, and here is how we limit it.” The approver should not be the PR author.
- A flag is not always a control. A default-off flag helps only if the off path is tested and someone owns turning it on. It does nothing for irreversible effects: a destructive migration behind a flag is still destructive once it runs.
Recording who approved what, and why, is the governance half of the problem that Two Halves of a Merge Gate describes. Without the record, an override is indistinguishable from the check never having run.
How do you enforce a severity policy in CI?
You enforce a severity policy in CI by making the required check fail only on the blocking set, and reporting everything else as a warning. The mistake to avoid is failing on the count of findings.
- name: Fail on reliability findings
run: |
count=$(jq '.findings | length' findings.json)
if [ "$count" -gt 0 ]; then exit 1; fi # old debt, S4s, and guesses all block
- name: Enforce reliability severity policy
run: |
# field names are illustrative; map them to your tool's output
blocking=$(jq '[.findings[]
| select(.status == "new" or .status == "activated")
| select(.verified == true)
| select(.severity == "S1" or (.severity == "S2" and .override == null))
] | length' findings.json)
warn=$(jq '[.findings[] | select(.severity == "S3")] | length' findings.json)
if [ "$warn" -gt 0 ]; then
echo "::warning::$warn S3 reliability findings need acknowledgement"
fi
if [ "$blocking" -gt 0 ]; then
echo "::error::$blocking blocking findings (S1, or S2 without an override)"
exit 1
fi
Make this step’s job a required status check on the protected branch, so a failure blocks the merge. The ::warning:: and ::error:: lines are GitHub Actions workflow commands that surface as annotations on the run. S1 overrides are deliberately not read from the file: they go through your platform’s admin bypass, which leaves its own audit trail.
How do you roll out a severity policy without slowing the team?
Roll out a severity policy in report-only mode first, then turn on blocking one level at a time, and only after the findings at that level have proven accurate. Teams that switch on blocking on day one usually switch it off by day thirty.
- Write the rubric with your own examples. Rate your last ten to twenty incidents with the impact, likelihood, blast radius, and reversibility table. Adjust the definitions until the ratings match what the team remembers.
- Run in report-only mode. For two to four weeks, compute what would have blocked but let everything merge. Post the would-be blocks as comments.
- Sample the would-be blocks. Have an engineer who knows the system check a sample of S1 and S2 findings against the evidence rule. Count how many were real.
- Block S1 only. Turn on the required check for verified S1 findings owned by the PR, with the emergency override path.
- Add S2 with overrides. Once S1 blocks are rarely disputed, add S2, with the override record described above.
- Review monthly and tune. Look at block, override, and confirmation rates, and change the rubric or the rules that produce disputed findings.
How do you know a merge gate policy is working?
A merge gate policy is working when it blocks rarely, the blocks are almost always confirmed real, overrides are occasional and followed up, and incidents are not traced back to changes that merged with warnings. Track these monthly.
| Signal | Healthy | If it drifts, it means |
|---|---|---|
| Block rate (PRs blocked / PRs merged) | Low and stable | Rising: severity inflation or a noisy rule. Near zero for months: check the gate still runs. |
| Confirmation rate of blocked findings | High | Falling: findings lack evidence, or a rule misfires in your codebase. Tune or demote it. |
| Override rate on S2 | Occasional | Rising: the S2 bar is too low, or teams are under delivery pressure the policy ignores. |
| Overdue override follow-ups | Few | Overrides are becoming permanent exceptions. Escalate or re-rate. |
| Time from block to resolution | Hours, not days | Findings are hard to act on. Improve suggested fixes and evidence. |
| Incidents from changes merged with S3 warnings | Rare | The rubric under-rates a class of finding. Raise it. |
The last row is the most valuable feedback loop. When an incident happens, check the change that caused it (see Which Commit Caused the Production Incident?) and ask what the policy said about it. A finding that was rated S3 and merged is a rubric bug, and fixing it is how the policy improves.
How Tomosu helps
Tomosu produces reliability findings for pull requests and the standing codebase, with the inputs a severity policy needs already attached, so the policy can be applied consistently rather than argued case by case:
- Evidence on each finding: the file and line, the call path that reaches it, and the condition that triggers it, so Question 2 can be answered by reading the finding.
- Ownership: because the repository is analyzed before the change, findings the PR introduces or activates are distinguished from pre-existing ones.
- Blast radius weighting: findings on critical paths and shared resources, such as connection pools and shared clients, are ranked above the same pattern in isolated code.
- A consistent signal: findings roll up into the Production Reliability Index, including Governance Compliance, which your team can use as an input to its own merge policy.
What blocks a merge remains your team’s decision. Tomosu’s job is to make sure that decision is made with evidence in front of the people making it.
Scan your repository with Tomosu →
Key takeaways
- Block only findings the PR introduced or activated, that are verified with evidence, and that rate S1 or S2.
- Rate severity from impact and likelihood, then raise it one level for a critical or shared path and one for irreversible effects.
- Pre-existing, unverified, and low-severity findings should never block on their own. They go to a backlog or a comment.
- Overrides are recorded decisions: a named approver who is not the author, a compensating control, a follow-up ticket, and an expiry.
- In CI, fail the required check on the blocking set only, never on the count of findings.
- Roll out in report-only mode, then block S1, then add S2, each step gated on confirmed accuracy.
- Measure block, confirmation, and override rates monthly, and treat incidents from warned changes as rubric bugs.
Frequently asked questions
Should a reliability finding block a merge?
Only when the pull request introduced or activated it, it is backed by evidence that cites the code and the trigger condition, and it is severe under an agreed framework: likely to cause data loss, an outage, or serious degradation on a critical path. Other findings should warn or be tracked, not block.
What is a severity framework for code review findings?
It is a shared rubric that rates each finding by impact, likelihood, blast radius, and reversibility, and maps each severity level to an action such as block, block unless overridden, warn, or track. It makes merge decisions consistent across reviewers and tools.
Should pre-existing issues block a pull request?
Not when the pull request does not change their exposure. Put them in a tracked backlog with an owner. If the pull request makes a pre-existing issue reachable, more frequent, or more severe, treat it as an activated finding and rate it like a new one.
What evidence should a blocking finding include?
The file and line, the condition that triggers the failure, the expected impact and the path it affects, a suggested fix, and how to verify the fix. A finding without that evidence should be a warning that asks for it, not a block.
How should merge gate overrides work?
Allow overrides for high but not critical findings when a named approver who is not the author records the reason, a compensating control such as a default-off feature flag, a canary, or a tested rollback, and a follow-up ticket with a due date. Review overrides monthly; a rising override rate means the policy or the findings need tuning.
How do you roll out a blocking reliability gate without slowing the team?
Start in report-only mode, measure how many findings would have blocked and how many of those were real, then block only the most severe level. Add the next level with an override path once those blocks are rarely disputed, and review the policy monthly.
What metrics show whether a merge gate policy is working?
Block rate, the share of blocked findings confirmed real, override rate, overdue override follow-ups, time from block to resolution, and incidents traced to changes that merged with warnings. A healthy gate blocks rarely and is almost always right when it does.
Can a feature flag replace fixing a severe finding?
Sometimes, as a compensating control for a high-severity finding, if the flag defaults off, the off path is tested, and someone owns turning it on. A flag does not help with irreversible effects such as a destructive migration or data already written, so it should not be used to override those.
A merge gate is only as good as the findings it blocks on. Tomosu puts the evidence, ownership, and blast radius on each finding so your policy can decide. Assess your repository →