Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Code review

AI Code Review False Positives: How to Verify a Reliability Finding

Tomosu AI·14 min read·

An AI reviewer comments on your pull request: “possible connection leak,” “missing timeout,” “race condition.” Some of those comments will prevent an incident. Some are wrong. Accept them all and you ship churn, sometimes risky churn. Dismiss them all and the real ones go with the noise. What you need is a quick, repeatable way to tell them apart.

Quick answer

A reliability finding is verified when you can state its trigger, confirm the code path is reachable in production, check that nothing outside the diff already prevents it, and produce evidence such as a failing test, a trace, or a documented default. Most AI code review false positives come from missing context, not bad logic.

This guide covers what counts as a false positive, why AI reviewers produce them on reliability issues in particular, a six-step verification procedure, the evidence that settles the most common finding types, and two worked examples: one false positive and one real finding that looked harmless.

What is a false positive in AI code review?

A false positive in AI code review is a finding that claims a problem the code does not actually have in the context where it runs. The claim may be plausible from the lines shown, but the behavior it predicts cannot happen, because of how the library works, how the code is called, or a guard that lives elsewhere.

That definition leaves out two outcomes that are often lumped in with false positives and need different handling:

FOUR OUTCOMES, FOUR DIFFERENT COSTS AI RAISED A FINDING AI STAYED SILENT PROBLEM IS REAL NO REAL PROBLEM True positive Fixed before merge. The reason to use a reviewer at all. False negative Ships. Found by users or on-call. The most expensive outcome. False positive Costs review time and trust. A needless “fix” can add new risk. True negative Nothing to do. TWO MORE OUTCOMES THAT ARE NOT FALSE POSITIVES True but low impact Real, rarely hit. Downgrade, don’t dismiss. Unverified Could be either. Needs evidence first.
A false positive is not free: it costs review time, trains people to skim findings, and sometimes produces a “fix” that makes the code worse.

The last point in the false positive cell is easy to miss. Accepting a wrong reliability finding is not neutral. “Add a retry here” can create retry amplification (see How to Spot a Retry Storm Before Merge). “Add a lock” can serialize a hot path. “Add a timeout” with a value picked at random can fail requests that were fine. Verifying before fixing protects the code from the reviewer as well as from the bug.

Why do AI code reviewers produce false positives on reliability issues?

Most reliability false positives happen because the facts that decide the finding are not in the diff. Whether a call can hang depends on client configuration. Whether a connection leaks depends on the framework’s lifecycle. Whether a race is possible depends on database constraints and how many workers run the code. An AI reviewer that sees only the changed lines, plus whatever context it retrieves, has to guess at those facts.

THE FACTS THAT DECIDE A FINDING USUALLY SIT OUTSIDE THE DIFF THE DIFF + resp = session.get(url) + db.insert(user) Shared client config Default timeouts, retries, adapters, interceptors Framework lifecycle Who opens and closes connections, transactions Data constraints Unique indexes, foreign keys, row locks Runtime config Workers, replicas, consumer concurrency, gateway timeouts Callers and entry points Who reaches this code, how often, with what data A finding that ignores these can be plausible and still wrong, or look harmless and still be real.
Verifying a reliability finding is mostly the work of pulling in the context that the reviewer did not have.

The table below lists the causes we see most often and the question that exposes each one.

Cause of the false positiveTypical findingQuestion that exposes it
Context outside the diff“No timeout on this HTTP call”Is a default set on the shared client, adapter, or interceptor this call uses?
Wrong API semantics“Connection is never closed”Does the framework or context manager close it? What do the library docs say?
Unreachable or non-production path“Unbounded loop over results”Is this code reached in production, or only in tests, migrations, or dead branches?
Invariant enforced elsewhere“Race: duplicate insert possible”Is there a unique constraint, a lock, or a single consumer that makes it impossible?
Scale assumptions“N+1 query in this loop”How large is the collection in production? Is it eager-loaded upstream?
Invented APIs or settings“Set maxRetries on this client”Does that option exist in this library and version?

The same gap runs the other way. Missing context also causes false negatives: a three-line change that calls a shared client can look safe in the diff and be risky because of the retries configured elsewhere. That is the central argument of Production Reliability vs Code Review, and it is why a repository-wide review finds things a PR diff misses.

Similar questions come up on Stack Overflow, usually in the form “the tool keeps flagging X, but X is handled.” The fix is rarely to argue with the reviewer. It is to make the handling visible, then record why the finding does not apply.

What does a verifiable reliability finding look like?

A verifiable finding states a trigger, a code path, an impact, and the evidence that would confirm or refute it. A vague finding names a pattern and a suggestion. The difference decides how long verification takes: minutes for the first, an open-ended investigation for the second.

SAME ISSUE, TWO FINDINGS VAGUE · HARD TO VERIFY “Possible resource issue in syncInventory. Consider limiting concurrency.” × no trigger × no impact × no way to confirm × no way to refute VERIFIABLE · FIVE FIELDS CLAIM TRIGGER CODE PATH IMPACT EVIDENCE One API call per SKU, all started at once Nightly job passes every SKU in a warehouse nightlySync → syncInventory → update Burst far above the rate limit: 429s, stock levels left stale Max SKUs per warehouse vs API limit. Refuted if calls are bounded upstream.
If a finding cannot be restated in these five fields, it is not ready to accept or dismiss. Ask for the missing field first.

You can apply this format whether the finding came from an AI reviewer, a static analyzer, or a colleague. If your AI review tool accepts custom instructions, asking it to include a trigger, an impact, and a “refuted if” condition in each reliability comment makes verification much faster, because it moves some of the work to the reviewer.

How do you verify an AI code review finding, step by step?

Work through these steps in order. Most false positives fall out at step 2, 3, or 4, usually within a few minutes. The steps that take longer, 5 in particular, are reserved for findings that survived the cheap checks.

  1. Restate the finding as a falsifiable claim. Write it as “under this trigger, this code path does this, causing this impact.” If you cannot, ask the reviewer or author for the missing part before doing anything else.
  2. Check the claim against the code and the library’s real semantics. Read the lines the finding cites and the functions they call. Check defaults and behavior in the library’s documentation for the version you use, not from memory.
  3. Check that the path is reachable in production. Find the callers and entry points. Rule out test-only code, dead branches, disabled flags, and one-off scripts, and note how often the path runs.
  4. Check context outside the diff. Look for shared client configuration, framework lifecycle, database constraints, locks, and runtime settings such as worker counts and replicas that already prevent or amplify the problem.
  5. Produce evidence. Write a failing test, reproduce it locally or in staging, pull a trace or a production query, or cite the documentation that settles it.
  6. Classify and record the decision. Mark it true positive, false positive, true but low impact, or unverified, with a one-line reason and a link to the evidence.
VERIFYING A RELIABILITY FINDING: FIVE QUESTIONS, IN ORDER Q1 · IS IT A CLAIM? Can it be restated as trigger, path, and impact? Q2 · CODE AND API SEMANTICS Does the code really behave as claimed? Q3 · REACHABILITY Is the path reached in production? Q4 · CONTEXT OUTSIDE THE DIFF Does something elsewhere already prevent it? Q5 · EVIDENCE Does a test, trace, or query reproduce it? Unverified Ask for the missing specifics False positive Cite the doc or line that refutes it False positive or low Unreachable, test-only, or rare False positive Cite the guard, with a link Unverified Time-box a test or ask for data NONONOYESNO YESYESYESNOYES True positive. Fix it, and set severity by the measured impact, not by how alarming the comment sounds.
Cheap checks first. Most AI code review false positives are settled by reading the library docs, the callers, or the shared configuration.
Verify the fix as well as the finding

When a finding is real, AI reviewers often attach a suggested fix. Run the same checks on it. A suggested timeout may be longer than the caller’s own deadline, a suggested retry may stack on retries configured elsewhere, and a suggested cache may never be invalidated. A true positive with a bad fix is still a risk.

What evidence confirms or refutes common reliability findings?

Each finding type has a cheapest piece of evidence that settles it. Reach for that first. A load test is rarely needed to dismiss a false positive; a link to the line that already handles the case usually is enough.

FindingConfirmed ifRefuted ifCheapest evidence
Missing timeoutNo timeout at the call, client, or adapter level, and the library default is unboundedA shared client, adapter, or interceptor sets one; or the call runs under an enforced deadlineThe client construction code, plus the library docs on defaults
Resource or connection leakA path (often an early return or exception) skips releaseContext manager, try-with-resources, or framework scope releases on every pathA test that raises mid-block and asserts the pool or handle count returns to baseline
Race or duplicate writeTwo workers can interleave read and write with no constraint or lockA unique constraint, row lock, or single-consumer guarantee prevents itThe schema or migration; a concurrent test that runs the path twice in parallel
N+1 queryQuery count grows with collection size on a production pathRelation is eager-loaded upstream, or collections are provably tinyQuery count assertion in a test with realistic data volume
Unsafe retryThe operation is not idempotent, or retries stack across layersAn idempotency key is sent and enforced; only one layer retriesThe server’s handling of the key; the client and gateway retry config
Unbounded concurrency or memoryInput size is unbounded or large in productionInput is capped by validation, pagination, or a batch sizeProduction size distribution (max, p99) for the input
Missing indexThe query plan shows a scan on a large tableAn existing index (including a composite prefix) covers the predicateEXPLAIN on a production-sized dataset

Some of these have their own deep dives in this series: duplicate writes in Why “Check Then Insert” Creates Duplicate Records, and query counts in N+1 Queries That Pass Tests but Fail Under Production Traffic.

Worked examples: one false positive, one real finding

Example 1: “Missing timeout” that is a false positive

A pull request adds a call to a pricing service. The AI reviewer comments: “session.get() is called without a timeout. requests waits indefinitely by default, so this call can hang the worker.”

billing/pricing.py · the diffflagged
from platform.http import session

def get_price(sku):
    resp = session.get(f"{PRICING_URL}/prices/{sku}")  # no timeout= here
    resp.raise_for_status()
    return resp.json()["amount"]

Step 2 half-confirms it: the requests documentation does say that requests do not time out unless you set a timeout explicitly. Step 4 is where it falls apart. The imported session is built in a shared module that mounts an adapter supplying a default timeout whenever the caller does not pass one:

platform/http.py · not in the diffthe guard
import requests
from requests.adapters import HTTPAdapter

class TimeoutHTTPAdapter(HTTPAdapter):
    def __init__(self, *args, timeout=(1, 3), **kwargs):
        self.timeout = timeout
        super().__init__(*args, **kwargs)

    def send(self, request, **kwargs):
        if kwargs.get("timeout") is None:
            kwargs["timeout"] = self.timeout   # (connect, read) seconds
        return super().send(request, **kwargs)

session = requests.Session()
session.mount("https://", TimeoutHTTPAdapter())
session.mount("http://", TimeoutHTTPAdapter())

The finding is a false positive for this call. The record should say exactly that, with the link: “Default (1 s, 3 s) timeout applied by TimeoutHTTPAdapter in platform/http.py.” Two follow-ups keep it honest. First, any code that calls requests.get() directly bypasses the adapter, so the same comment on such a call would be a true positive. Second, if the guard is invisible to reviewers, a short comment at the call site (“timeout from shared session”) prevents the same false positive on the next pull request.

Example 2: “Unbounded concurrency” that looks harmless and is real

A different pull request refactors an inventory sync to run in parallel. The AI reviewer’s comment is short and easy to wave away: “Consider limiting concurrency here.”

jobs/sync-inventory.ts · the diffreal finding
export async function syncInventory(skus: string[]) {
  // Every update starts immediately; Promise.all only waits for them.
  await Promise.all(skus.map((sku) => inventoryApi.update(sku)));
}

Restated (step 1): under the nightly sync, syncInventory starts one request per SKU at the same time, which can exceed the inventory API’s rate limit and leave stock levels stale. Step 3 finds the caller: the nightly job passes every SKU in a warehouse. Step 5 checks the input size in production: the largest warehouse in this hypothetical has tens of thousands of SKUs, against a rate limit measured in tens of requests per second. The finding is a true positive, and the one-line comment undersold it.

jobs/sync-inventory.ts · fixbounded
import pLimit from "p-limit";

const limit = pLimit(10);  // at most 10 updates in flight

export async function syncInventory(skus: string[]) {
  await Promise.all(skus.map((sku) => limit(() => inventoryApi.update(sku))));
}

The p-limit library caps how many of the wrapped calls run at once. Choose the limit from the dependency’s documented rate limit and latency, and handle 429 responses with backoff rather than failing the whole batch. The same code in a CLI script that only ever processes a handful of items would be true but low impact: worth a note, not a blocker.

The severity of a finding lives in the callers and the data, not in the comment. A quiet comment can be the real one.

How do you record a dismissal so the same false positive does not return?

Record every dismissal with a category, a one-line reason, and a link to the evidence, and scope any suppression as narrowly as possible. The record helps the next reviewer, gives you data about which kinds of findings are noisy, and makes a later incident review possible if the dismissal turns out to be wrong.

PR comment · decision recordcopy this
# Decision is one of: true positive | false positive | low impact | unverified
Decision: false positive
Finding:  missing timeout on session.get() in billing/pricing.py
Reason:   default (1 s, 3 s) timeout applied by TimeoutHTTPAdapter
Evidence: platform/http.py L4-L12; requests docs on timeouts
Scope:    this call only; direct requests.get() calls are still in scope

What should you do with a finding you cannot verify?

Treat an unverified finding as an open question with an owner and a deadline, not as a pass or a block by default. What to do next depends on what it would cost if the finding were true.

If the finding were true…Reasonable defaultEvidence to ask for
Data loss, duplicate charges, or corruptionHold the merge until verifiedA concurrent test or a documented guarantee (constraint, idempotency key)
Outage on a user-facing pathHold, or merge behind a flag with a staged rolloutLoad test at production volume, or production size data for the input
Degraded latency on a secondary pathMerge with a tracked follow-upA metric or alert that would show it within a day
Internal tooling or batch job with retriesMerge and note itNone required; revisit if it recurs

This is a starting point, not a policy. Whether a finding should block a merge is a governance decision that depends on severity, confidence, and blast radius together; Should a Reliability Finding Block a Merge? A Severity Framework covers that in depth, and A Signal Is Not a Gate explains why a raw finding should rarely be a gate on its own.

Do not let “AI said so” be the evidence

Asking a second AI tool, or the same one again, whether a finding is valid is not verification. It is another opinion with the same missing context. Evidence is something you can link to: a line of code, a documented default, a test result, a trace, or a production query.

How Tomosu helps

Most reliability false positives, and many false negatives, come from context that sits outside the diff. Tomosu analyzes the repository as a whole and each change against it, which is the context the verification steps above ask for:

These signals roll up into the Production Reliability Index. Tomosu is one input to the review, not a replacement for the checks above; its findings should be verified the same way.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

What is a false positive in AI code review?

It is a finding that claims a problem the code does not actually have in the context where it runs. The claim may look plausible from the changed lines, but the predicted behavior cannot happen because of how the library works, how the code is called, or a guard that lives outside the diff, such as a shared client default or a database constraint.

Why do AI code reviewers produce false positives on reliability issues?

Because the facts that decide most reliability findings are not in the diff. Whether a call can hang, a connection can leak, or a race can occur depends on client configuration, framework lifecycle, database constraints, callers, and runtime settings such as worker counts. A reviewer without that context has to guess, and some guesses are wrong.

How do I verify an AI code review finding about reliability?

Restate it as a claim with a trigger, a code path, and an impact. Check the code against the library’s documented behavior, confirm the path is reachable in production, and look for guards outside the diff. Then produce evidence, such as a failing test, a trace, or a production query, and record the decision with a link to that evidence.

What counts as evidence for or against a reliability finding?

Something you can link to: the line of code or configuration that handles the case, the library documentation for the version you use, a test that fails before a fix and passes after, a trace or log from production, a query plan, or production data such as the largest input size. A second AI opinion is not evidence.

Should I dismiss an AI finding I cannot verify?

Not automatically. Treat it as an open question with an owner and a deadline. If the finding would mean data loss or an outage on a user-facing path if true, hold the merge or ship behind a staged rollout until it is verified. For low-impact paths, merge with a tracked follow-up.

Is a finding in test-only or unreachable code a false positive?

For production reliability, usually yes, because the predicted failure cannot reach users. Record it as a false positive or low impact with the reason. Check reachability carefully, though: code that looks unused may be called through reflection, a job scheduler, a feature flag, or a message consumer.

Can accepting a false positive make code less reliable?

Yes. Fixes suggested for findings that are not real can add risk: a retry that stacks with retries elsewhere, a lock that serializes a hot path, a cache with no invalidation, or a timeout shorter than normal latency. Verify the suggested fix with the same checks as the finding.

How can I reduce AI code review false positives over time?

Record each decision by finding category and look for the categories that are mostly wrong. Make guards visible near the code they protect, give the review tool more repository context or custom instructions where it supports them, and ask for findings that state a trigger, an impact, and what would refute them. Suppress individual instances rather than whole rules.


An AI reviewer is most useful when its findings can be checked quickly. Tomosu brings the repository context that makes that possible: callers, configuration, and blast radius, alongside the diff. Assess your repository →