An AI reviewer comments on your pull request: “possible connection leak,” “missing timeout,” “race condition.” Some of those comments will prevent an incident. Some are wrong. Accept them all and you ship churn, sometimes risky churn. Dismiss them all and the real ones go with the noise. What you need is a quick, repeatable way to tell them apart.
A reliability finding is verified when you can state its trigger, confirm the code path is reachable in production, check that nothing outside the diff already prevents it, and produce evidence such as a failing test, a trace, or a documented default. Most AI code review false positives come from missing context, not bad logic.
- Restate the finding as “under X, this path does Y, causing Z.”
- Check the code and the library’s actual semantics.
- Check reachability and context outside the diff.
- Get evidence, then classify and record the decision.
This guide covers what counts as a false positive, why AI reviewers produce them on reliability issues in particular, a six-step verification procedure, the evidence that settles the most common finding types, and two worked examples: one false positive and one real finding that looked harmless.
What is a false positive in AI code review?
A false positive in AI code review is a finding that claims a problem the code does not actually have in the context where it runs. The claim may be plausible from the lines shown, but the behavior it predicts cannot happen, because of how the library works, how the code is called, or a guard that lives elsewhere.
That definition leaves out two outcomes that are often lumped in with false positives and need different handling:
- True but low impact. The problem is real, but the path is rarely used, small, or not user-facing. For example, unbounded concurrency in a one-off admin script that processes 20 rows. It deserves a lower severity, not a dismissal.
- Unverified. Nobody can tell yet. The finding may be right, but confirming it needs data you do not have in review, such as production volumes or a dependency’s rate limit.
The last point in the false positive cell is easy to miss. Accepting a wrong reliability finding is not neutral. “Add a retry here” can create retry amplification (see How to Spot a Retry Storm Before Merge). “Add a lock” can serialize a hot path. “Add a timeout” with a value picked at random can fail requests that were fine. Verifying before fixing protects the code from the reviewer as well as from the bug.
Why do AI code reviewers produce false positives on reliability issues?
Most reliability false positives happen because the facts that decide the finding are not in the diff. Whether a call can hang depends on client configuration. Whether a connection leaks depends on the framework’s lifecycle. Whether a race is possible depends on database constraints and how many workers run the code. An AI reviewer that sees only the changed lines, plus whatever context it retrieves, has to guess at those facts.
The table below lists the causes we see most often and the question that exposes each one.
| Cause of the false positive | Typical finding | Question that exposes it |
|---|---|---|
| Context outside the diff | “No timeout on this HTTP call” | Is a default set on the shared client, adapter, or interceptor this call uses? |
| Wrong API semantics | “Connection is never closed” | Does the framework or context manager close it? What do the library docs say? |
| Unreachable or non-production path | “Unbounded loop over results” | Is this code reached in production, or only in tests, migrations, or dead branches? |
| Invariant enforced elsewhere | “Race: duplicate insert possible” | Is there a unique constraint, a lock, or a single consumer that makes it impossible? |
| Scale assumptions | “N+1 query in this loop” | How large is the collection in production? Is it eager-loaded upstream? |
| Invented APIs or settings | “Set maxRetries on this client” | Does that option exist in this library and version? |
The same gap runs the other way. Missing context also causes false negatives: a three-line change that calls a shared client can look safe in the diff and be risky because of the retries configured elsewhere. That is the central argument of Production Reliability vs Code Review, and it is why a repository-wide review finds things a PR diff misses.
Similar questions come up on Stack Overflow, usually in the form “the tool keeps flagging X, but X is handled.” The fix is rarely to argue with the reviewer. It is to make the handling visible, then record why the finding does not apply.
What does a verifiable reliability finding look like?
A verifiable finding states a trigger, a code path, an impact, and the evidence that would confirm or refute it. A vague finding names a pattern and a suggestion. The difference decides how long verification takes: minutes for the first, an open-ended investigation for the second.
You can apply this format whether the finding came from an AI reviewer, a static analyzer, or a colleague. If your AI review tool accepts custom instructions, asking it to include a trigger, an impact, and a “refuted if” condition in each reliability comment makes verification much faster, because it moves some of the work to the reviewer.
How do you verify an AI code review finding, step by step?
Work through these steps in order. Most false positives fall out at step 2, 3, or 4, usually within a few minutes. The steps that take longer, 5 in particular, are reserved for findings that survived the cheap checks.
- Restate the finding as a falsifiable claim. Write it as “under this trigger, this code path does this, causing this impact.” If you cannot, ask the reviewer or author for the missing part before doing anything else.
- Check the claim against the code and the library’s real semantics. Read the lines the finding cites and the functions they call. Check defaults and behavior in the library’s documentation for the version you use, not from memory.
- Check that the path is reachable in production. Find the callers and entry points. Rule out test-only code, dead branches, disabled flags, and one-off scripts, and note how often the path runs.
- Check context outside the diff. Look for shared client configuration, framework lifecycle, database constraints, locks, and runtime settings such as worker counts and replicas that already prevent or amplify the problem.
- Produce evidence. Write a failing test, reproduce it locally or in staging, pull a trace or a production query, or cite the documentation that settles it.
- Classify and record the decision. Mark it true positive, false positive, true but low impact, or unverified, with a one-line reason and a link to the evidence.
When a finding is real, AI reviewers often attach a suggested fix. Run the same checks on it. A suggested timeout may be longer than the caller’s own deadline, a suggested retry may stack on retries configured elsewhere, and a suggested cache may never be invalidated. A true positive with a bad fix is still a risk.
What evidence confirms or refutes common reliability findings?
Each finding type has a cheapest piece of evidence that settles it. Reach for that first. A load test is rarely needed to dismiss a false positive; a link to the line that already handles the case usually is enough.
| Finding | Confirmed if | Refuted if | Cheapest evidence |
|---|---|---|---|
| Missing timeout | No timeout at the call, client, or adapter level, and the library default is unbounded | A shared client, adapter, or interceptor sets one; or the call runs under an enforced deadline | The client construction code, plus the library docs on defaults |
| Resource or connection leak | A path (often an early return or exception) skips release | Context manager, try-with-resources, or framework scope releases on every path | A test that raises mid-block and asserts the pool or handle count returns to baseline |
| Race or duplicate write | Two workers can interleave read and write with no constraint or lock | A unique constraint, row lock, or single-consumer guarantee prevents it | The schema or migration; a concurrent test that runs the path twice in parallel |
| N+1 query | Query count grows with collection size on a production path | Relation is eager-loaded upstream, or collections are provably tiny | Query count assertion in a test with realistic data volume |
| Unsafe retry | The operation is not idempotent, or retries stack across layers | An idempotency key is sent and enforced; only one layer retries | The server’s handling of the key; the client and gateway retry config |
| Unbounded concurrency or memory | Input size is unbounded or large in production | Input is capped by validation, pagination, or a batch size | Production size distribution (max, p99) for the input |
| Missing index | The query plan shows a scan on a large table | An existing index (including a composite prefix) covers the predicate | EXPLAIN on a production-sized dataset |
Some of these have their own deep dives in this series: duplicate writes in Why “Check Then Insert” Creates Duplicate Records, and query counts in N+1 Queries That Pass Tests but Fail Under Production Traffic.
Worked examples: one false positive, one real finding
Example 1: “Missing timeout” that is a false positive
A pull request adds a call to a pricing service. The AI reviewer comments: “session.get() is called without a timeout. requests waits indefinitely by default, so this call can hang the worker.”
from platform.http import session
def get_price(sku):
resp = session.get(f"{PRICING_URL}/prices/{sku}") # no timeout= here
resp.raise_for_status()
return resp.json()["amount"]
Step 2 half-confirms it: the requests documentation does say that requests do not time out unless you set a timeout explicitly. Step 4 is where it falls apart. The imported session is built in a shared module that mounts an adapter supplying a default timeout whenever the caller does not pass one:
import requests
from requests.adapters import HTTPAdapter
class TimeoutHTTPAdapter(HTTPAdapter):
def __init__(self, *args, timeout=(1, 3), **kwargs):
self.timeout = timeout
super().__init__(*args, **kwargs)
def send(self, request, **kwargs):
if kwargs.get("timeout") is None:
kwargs["timeout"] = self.timeout # (connect, read) seconds
return super().send(request, **kwargs)
session = requests.Session()
session.mount("https://", TimeoutHTTPAdapter())
session.mount("http://", TimeoutHTTPAdapter())
The finding is a false positive for this call. The record should say exactly that, with the link: “Default (1 s, 3 s) timeout applied by TimeoutHTTPAdapter in platform/http.py.” Two follow-ups keep it honest. First, any code that calls requests.get() directly bypasses the adapter, so the same comment on such a call would be a true positive. Second, if the guard is invisible to reviewers, a short comment at the call site (“timeout from shared session”) prevents the same false positive on the next pull request.
Example 2: “Unbounded concurrency” that looks harmless and is real
A different pull request refactors an inventory sync to run in parallel. The AI reviewer’s comment is short and easy to wave away: “Consider limiting concurrency here.”
export async function syncInventory(skus: string[]) {
// Every update starts immediately; Promise.all only waits for them.
await Promise.all(skus.map((sku) => inventoryApi.update(sku)));
}
Restated (step 1): under the nightly sync, syncInventory starts one request per SKU at the same time, which can exceed the inventory API’s rate limit and leave stock levels stale. Step 3 finds the caller: the nightly job passes every SKU in a warehouse. Step 5 checks the input size in production: the largest warehouse in this hypothetical has tens of thousands of SKUs, against a rate limit measured in tens of requests per second. The finding is a true positive, and the one-line comment undersold it.
import pLimit from "p-limit";
const limit = pLimit(10); // at most 10 updates in flight
export async function syncInventory(skus: string[]) {
await Promise.all(skus.map((sku) => limit(() => inventoryApi.update(sku))));
}
The p-limit library caps how many of the wrapped calls run at once. Choose the limit from the dependency’s documented rate limit and latency, and handle 429 responses with backoff rather than failing the whole batch. The same code in a CLI script that only ever processes a handful of items would be true but low impact: worth a note, not a blocker.
The severity of a finding lives in the callers and the data, not in the comment. A quiet comment can be the real one.
How do you record a dismissal so the same false positive does not return?
Record every dismissal with a category, a one-line reason, and a link to the evidence, and scope any suppression as narrowly as possible. The record helps the next reviewer, gives you data about which kinds of findings are noisy, and makes a later incident review possible if the dismissal turns out to be wrong.
# Decision is one of: true positive | false positive | low impact | unverified
Decision: false positive
Finding: missing timeout on session.get() in billing/pricing.py
Reason: default (1 s, 3 s) timeout applied by TimeoutHTTPAdapter
Evidence: platform/http.py L4-L12; requests docs on timeouts
Scope: this call only; direct requests.get() calls are still in scope
- Suppress narrowly. Dismiss the instance, not the whole rule. A rule that is wrong once is often right elsewhere, as Example 1 shows.
- Fix the context, not only the comment. If the guard lives far from the call, a short code comment, a typed wrapper, or a clearer name removes the ambiguity for humans and tools alike.
- Count by category. Tally decisions by finding type over a few weeks. A type with mostly false positives is a candidate for tuning or custom instructions; a type with mostly true positives deserves more weight in review.
- Revisit dismissals after incidents. When an RCA names a code path, check whether a dismissed finding pointed at it. See How to Write an RCA That Connects the Trigger, Code Path, and User Impact.
What should you do with a finding you cannot verify?
Treat an unverified finding as an open question with an owner and a deadline, not as a pass or a block by default. What to do next depends on what it would cost if the finding were true.
| If the finding were true… | Reasonable default | Evidence to ask for |
|---|---|---|
| Data loss, duplicate charges, or corruption | Hold the merge until verified | A concurrent test or a documented guarantee (constraint, idempotency key) |
| Outage on a user-facing path | Hold, or merge behind a flag with a staged rollout | Load test at production volume, or production size data for the input |
| Degraded latency on a secondary path | Merge with a tracked follow-up | A metric or alert that would show it within a day |
| Internal tooling or batch job with retries | Merge and note it | None required; revisit if it recurs |
This is a starting point, not a policy. Whether a finding should block a merge is a governance decision that depends on severity, confidence, and blast radius together; Should a Reliability Finding Block a Merge? A Severity Framework covers that in depth, and A Signal Is Not a Gate explains why a raw finding should rarely be a gate on its own.
Asking a second AI tool, or the same one again, whether a finding is valid is not verification. It is another opinion with the same missing context. Evidence is something you can link to: a line of code, a documented default, a test result, a trace, or a production query.
How Tomosu helps
Most reliability false positives, and many false negatives, come from context that sits outside the diff. Tomosu analyzes the repository as a whole and each change against it, which is the context the verification steps above ask for:
- Call paths and entry points: which endpoints, jobs, and consumers reach a changed function, which answers the reachability question in step 3.
- Surrounding configuration: shared clients, timeouts, retries, and transaction boundaries on the path, which is where most “already handled” guards live.
- Blast radius: which services and user flows depend on the changed code, so a finding on a checkout path is weighed differently from one in a nightly report.
- Findings framed as claims: the risk, the code path it runs through, and the evidence needed to confirm it, so reviewers can check a finding rather than take it on trust.
These signals roll up into the Production Reliability Index. Tomosu is one input to the review, not a replacement for the checks above; its findings should be verified the same way.
Scan your repository with Tomosu →
Key takeaways
- A false positive claims a problem the code cannot have in its real context. “True but low impact” and “unverified” are different outcomes with different handling.
- Most AI code review false positives on reliability come from missing context: shared client config, framework lifecycle, data constraints, callers, and runtime settings.
- Restate every finding as trigger, code path, and impact. If it cannot be restated, ask for specifics before acting.
- Run the cheap checks first: library semantics, reachability, and guards outside the diff settle most findings in minutes.
- Evidence is something you can link to. Another AI opinion is not evidence.
- Verify suggested fixes too. A wrong retry, lock, or timeout can make the code less reliable.
- Record decisions with a category, reason, and evidence link, and suppress narrowly.
Frequently asked questions
What is a false positive in AI code review?
It is a finding that claims a problem the code does not actually have in the context where it runs. The claim may look plausible from the changed lines, but the predicted behavior cannot happen because of how the library works, how the code is called, or a guard that lives outside the diff, such as a shared client default or a database constraint.
Why do AI code reviewers produce false positives on reliability issues?
Because the facts that decide most reliability findings are not in the diff. Whether a call can hang, a connection can leak, or a race can occur depends on client configuration, framework lifecycle, database constraints, callers, and runtime settings such as worker counts. A reviewer without that context has to guess, and some guesses are wrong.
How do I verify an AI code review finding about reliability?
Restate it as a claim with a trigger, a code path, and an impact. Check the code against the library’s documented behavior, confirm the path is reachable in production, and look for guards outside the diff. Then produce evidence, such as a failing test, a trace, or a production query, and record the decision with a link to that evidence.
What counts as evidence for or against a reliability finding?
Something you can link to: the line of code or configuration that handles the case, the library documentation for the version you use, a test that fails before a fix and passes after, a trace or log from production, a query plan, or production data such as the largest input size. A second AI opinion is not evidence.
Should I dismiss an AI finding I cannot verify?
Not automatically. Treat it as an open question with an owner and a deadline. If the finding would mean data loss or an outage on a user-facing path if true, hold the merge or ship behind a staged rollout until it is verified. For low-impact paths, merge with a tracked follow-up.
Is a finding in test-only or unreachable code a false positive?
For production reliability, usually yes, because the predicted failure cannot reach users. Record it as a false positive or low impact with the reason. Check reachability carefully, though: code that looks unused may be called through reflection, a job scheduler, a feature flag, or a message consumer.
Can accepting a false positive make code less reliable?
Yes. Fixes suggested for findings that are not real can add risk: a retry that stacks with retries elsewhere, a lock that serializes a hot path, a cache with no invalidation, or a timeout shorter than normal latency. Verify the suggested fix with the same checks as the finding.
How can I reduce AI code review false positives over time?
Record each decision by finding category and look for the categories that are mostly wrong. Make guards visible near the code they protect, give the review tool more repository context or custom instructions where it supports them, and ask for findings that state a trigger, an impact, and what would refute them. Suppress individual instances rather than whole rules.
An AI reviewer is most useful when its findings can be checked quickly. Tomosu brings the repository context that makes that possible: callers, configuration, and blast radius, alongside the diff. Assess your repository →