Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Code review

What a Repository-Wide Review Finds That a PR Diff Misses

Tomosu AI·13 min read·

A pull request diff is a list of changed lines. Whether those lines are safe usually depends on lines that did not change: the caller that holds a transaction, the shared client that retries, the job that calls the function fifty thousand times, the three other files that copy the same pattern. A reviewer who reads only the diff is judging a change without the code it runs inside.

Quick answer

A repository-wide review evaluates a change against the whole codebase, not just the changed hunks. It finds risks whose evidence lives in other files, which a PR diff cannot show even when every changed line is correct.

This is not an argument against reading diffs. Diff review is where intent, correctness, and design get judged, and nothing replaces it. The point is narrower: some of the most expensive production failures come from changes that look correct in the diff, because the reason they fail is in a file the diff never shows. This guide covers what those blind spots are, how to find them by hand, and how to classify what you find so it does not turn every pull request into an audit.

What does a PR diff actually show a reviewer?

A PR diff shows the lines that were added, removed, or modified, plus a few lines of surrounding context. On GitHub and GitLab the default view is the changed files, with the option to expand hidden lines inside those files. Files the pull request did not touch do not appear at all, however much they depend on the change.

That is a sensible design for reading a change. It is a poor design for judging a change’s effect on production, because production behavior is a property of the call path, not of the file. A function’s cost depends on who calls it, how often, and while holding what. A new network call’s worst case depends on the client it goes through. A validation rule’s strength depends on whether every write path goes through it.

THE DIFF WINDOW VS THE CODE IT RUNS INSIDE DIFF WINDOW · +2 −1 get_price(sku) now calls tax API, timeout=2 orders/checkout.py calls it with rows locked config/database.py pool_size=10, shared platform/http.py shared session, 2 retries jobs/reprice_catalog.py loops over 50,000 SKUs gateway/checkout-route upstream timeout 3 s tests/test_pricing.py mocks the tax API Everything outside the dashed box decides whether the change is safe. None of it is in the diff.
The diff is three lines in one file. The six files around it are what a repository-wide review reads to decide whether those three lines are safe.

A repository-wide review is a review that evaluates a change against the whole codebase: its callers and callees, the configuration it inherits, other code that follows the same pattern, and the invariants enforced elsewhere. It can also scan code the pull request did not touch, which matters when the change makes old code newly reachable.

How can a clean diff still break production?

Here is a complete pull request. A pricing function that used to read from an in-memory table now adds tax from a tax service. The author added a timeout, which is more than many reviewers would ask for. The tests pass.

pricing/service.py · the whole difflooks fine
 def get_price(sku: str) -> Decimal:
     base = _price_table[sku]
-    return base
+    rate = http.get(f"{TAX_URL}/rates/{sku}", timeout=2).json()["rate"]
+    return base * (1 + Decimal(str(rate)))

Nothing in those lines is wrong. Now read two files that are not in the diff. The first is the shared HTTP session the new call goes through:

platform/http.py · not in the diff
from requests import Session
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

http = Session()
# GETs retried on connect/read errors and 502/503/504
http.mount("https://", HTTPAdapter(max_retries=Retry(
    total=2, backoff_factor=0.5, status_forcelist=[502, 503, 504])))

The second is the checkout path that calls get_price:

orders/checkout.py · not in the diffnow unsafe
def place_order(session: Session, cart: Cart) -> Order:
    with session.begin():
        # connection held and stock rows locked from here...
        stock = session.scalars(
            select(Stock).where(Stock.sku.in_(cart.skus)).with_for_update()
        ).all()
        reserve(stock, cart)
        lines = [
            OrderLine(sku=i.sku, qty=i.qty,
                      unit_price=get_price(i.sku))  # ...to here
            for i in cart.items
        ]
        order = Order(customer_id=cart.customer_id, lines=lines)
        session.add(order)
    return order

Put together, the change means every checkout holds a database connection and row locks on popular SKUs while it makes one tax call per cart item, each of which can be attempted three times with backoff. A healthy tax service adds tens of milliseconds per item. A slow one turns a 40 ms transaction into many seconds. Ten concurrent slow checkouts pin a ten-connection pool, and every other request that needs the database waits. Meanwhile the gateway gives up after 3 seconds, so the client may retry an order the server is still processing.

ONE CHECKOUT TRANSACTION, BEFORE AND AFTER BEFORE THE PR tx · ~40 ms lock rows, read 3 prices from memory, insert, commit AFTER THE PR, TAX SERVICE SLOW lock tax call 1 · up to 3 tries tax call 2 · up to 3 tries tax call 3 · ... commit Connection and row locks held for the whole bar: seconds, not milliseconds. GATEWAY · 3 S UPSTREAM TIMEOUT Client sees an error; server may still commit POOL · 10 CONNECTIONS 10 slow checkouts block every DB request
The diff added a network call. The caller turned it into a network call made while holding a connection and row locks, and the shared session multiplied its worst case.

None of that is visible from pricing/service.py. The reviewer would have had to know that place_order calls get_price inside session.begin(), that the session retries, and that a nightly job calls the same function for every SKU in the catalog, turning one deploy into 50,000 or more tax calls per night. And the unit test mocks http.get, so it proves the arithmetic and nothing about latency. That is the pattern behind many PRs that pass CI but break production.

The fix is not in the diff’s file either. It is a design change at the caller: fetch prices in one batched call before the transaction opens, and give the batch job its own rate-limited path.

orders/checkout.py · fixedfix
def place_order(session: Session, cart: Cart) -> Order:
    # one batched call, before any connection or lock is held
    prices = get_prices(cart.skus)
    with session.begin():
        stock = session.scalars(
            select(Stock).where(Stock.sku.in_(cart.skus)).with_for_update()
        ).all()
        reserve(stock, cart)
        lines = [OrderLine(sku=i.sku, qty=i.qty, unit_price=prices[i.sku])
                 for i in cart.items]
        order = Order(customer_id=cart.customer_id, lines=lines)
        session.add(order)
    return order

The diff tells you what changed. The repository tells you what the change will do.

What does a repository-wide review find that a PR diff misses?

A repository-wide review finds six kinds of risk that a PR diff cannot show, because their evidence lives in other files. Each has a different place to look.

1. Caller context: loops, transactions, locks, and hot paths

A changed function inherits its callers’ circumstances. A new database query is harmless in a settings page and an N+1 problem inside a loop. A new remote call is fine in a background worker and a pool-exhaustion risk inside a transaction. The question to answer is not “is this function correct?” but “who calls this, how often, and while holding what?” For the loop case in particular, see N+1 queries that pass tests but fail under production traffic.

2. Inherited configuration: retries, timeouts, pools, and flags

New code rarely configures its own client. It borrows a shared HTTP session, database pool, message producer, or cache client, and it inherits that object’s retry policy, timeouts, and limits. A diff that adds retries=3 looks modest until you find that the shared client already retries and the gateway retries too. That multiplication is the subject of How to Spot a Retry Storm Before Merge. The same applies to timeouts that no longer fit inside the caller’s deadline, and to feature flags whose default lives in a config file the PR did not touch.

3. Contract drift: return values, exceptions, and side effects

When a function starts returning None instead of raising, raising a new exception type, returning a lazily evaluated iterator instead of a list, or writing to a table it used to only read, every caller that relied on the old behavior is affected. The diff shows the new behavior. It does not show the caller that catches KeyError and now never sees it, or the caller that iterates the result twice.

4. Sibling copies: the same pattern in other files

Codebases repeat themselves. When a PR fixes a bug in one place, the same bug often exists in the files that were copied from it. When a PR introduces a pattern, a reviewer should know whether the repository already has a safer established version of it.

A FIX IN THE DIFF, THREE COPIES OUTSIDE IT IN THE DIFF · FIXED webhooks/stripe.py check-then-insert → upsert SEARCH THE REPO same pattern? webhooks/paypal.py same race, unchanged webhooks/adyen.py same race, unchanged jobs/replay_events.py same race, unchanged Diff review: duplicate bug fixed Repo-wide review: 1 of 4 copies fixed
A diff can only show the copy it fixed. Finding the other copies takes a search across the repository for the same pattern.

The webhook race in this figure is the check-then-insert pattern covered in Why “Check Then Insert” Creates Duplicate Records. It is a good example because the fix, a unique constraint plus an insert that handles conflicts, is short and local, which makes it easy to apply to one handler and forget the rest.

5. Bypassed invariants: rules enforced somewhere else

Many correctness rules are enforced in exactly one place: a service method that validates input, a repository class that adds the tenant filter, a wrapper that takes a lock or checks an idempotency key. A PR that adds a new write path, for example a bulk import endpoint that writes directly with the ORM, can skip that place without any line in the diff looking wrong. You only see the bypass by knowing where the rule is enforced and checking whether the new path goes through it.

6. Activated standing risk: old code the change makes reachable

Some risk has been sitting in the repository for years without causing an incident, because nothing important called it. A client with no timeout that was used by a monthly report. A query without an index on a table that used to be small. When a PR starts calling that code from a request path, the old risk becomes the new change’s risk, even though none of its lines are in the diff. This is the category most specific to repository-wide review, and the reason a standing scan of the codebase is worth having before any PR arrives.

Blind spotWhat the diff showsWhere the evidence lives
Caller contextA new query or remote call in a functionEvery call site: loops, transactions, locks, request versus batch paths
Inherited configurationA call through an existing clientWhere the client, pool, retry policy, timeout, or flag default is defined, including deploy manifests
Contract driftA changed return value, exception, or side effectCallers that depend on the old behavior: except blocks, null checks, repeated iteration
Sibling copiesA fix or a new pattern in one fileOther files with the same code shape, often copied from each other
Bypassed invariantsA new write or read pathThe single place that validates, filters by tenant, locks, or deduplicates
Activated standing riskA new call to existing codeThe callee’s own unguarded calls, missing indexes, and unbounded reads

How should findings outside the diff be classified?

Findings outside the diff should be split into three groups: new, activated, and pre-existing. The split matters because a repository-wide review will find things, and if every finding is treated as the PR author’s problem, reviews stall and people stop trusting the tool.

WHOSE FINDING IS IT? INPUT Baseline repo scan INPUT The pull request Q1 Did the PR introduce it? Q2 Does the PR make it reachable, more frequent, or worse? YESNOYESNO New Review in this PR Activated Also this PR’s risk Pre-existing Tracked backlog, no block
A baseline scan of the repository is what makes the split possible. Without it, every old finding looks new the first time someone reviews nearby code.

Which of the new and activated findings should actually stop a merge is a separate policy question, covered in Should a Reliability Finding Block a Merge? A Severity Framework.

A cross-file finding must cite the other file

A finding about a caller, a shared client, or a sibling copy is only checkable if it names the file and line it is based on. “This may be called in a loop” is a guess. “jobs/reprice_catalog.py:41 calls this once per SKU” is evidence. For how to check a finding before acting on it, see AI Code Review False Positives: How to Verify a Reliability Finding.

How do you review a pull request with repository context?

You review a pull request with repository context by starting from the changed symbols, not the changed files, and following each one outward. You do not need to do this for every PR. Do it for changes to shared code, hot paths, clients, data access, and anything that adds I/O.

  1. List the changed symbols. Write down the functions, classes, config keys, schemas, and message formats whose behavior changed, not just the files.
  2. Find every caller and consumer. Use find usages or a call hierarchy in your IDE, or git grep. For each call site, note whether it runs in a loop, inside a transaction or lock, on a request path, or in a batch job.
  3. Trace the configuration the change inherits. Find where the client, pool, retry policy, timeout, and feature flag default it uses are defined, including gateway and deployment files in the repository.
  4. Search for sibling copies. Search for the same code shape the PR fixes or introduces, and check whether the repository already has an established safer version.
  5. Check invariants enforced elsewhere. Identify where validation, tenant filtering, locking, and idempotency are enforced, and confirm the new path goes through them.
  6. Classify what you find. Mark each finding new, activated, or pre-existing, and keep pre-existing findings out of this PR’s review thread.

A few commands cover most of the manual work in a Git repository:

shell · repository context for a review
# function and class definitions touched by this branch
git diff --unified=0 origin/main...HEAD -- '*.py' \
  | grep -E '^[+-][[:space:]]*(def|class) '

# every call site of a changed function, across the whole repository
git grep -n -w 'get_price' -- '*.py'

# where shared clients configure retries and timeouts
git grep -n -E 'Retry\(|max_retries|timeout=' -- '*.py' '*.yaml'

# sibling copies of the pattern being fixed
git grep -n -E '\.filter_by\(event_id=' -- 'webhooks/' 'jobs/'

Text search is a floor, not a ceiling. It misses callers wired by name: decorators, dependency injection, message handlers subscribed by topic, scheduled jobs configured in YAML, and reflection. Check those registries separately. Static call graphs have the same gap for dynamic dispatch. Developers ask about this regularly on Stack Overflow, and the honest answer is that no single tool finds every caller.

QuestionDiff-only reviewRepository-wide review
Is the changed code correct?Yes, this is what it is forSame, plus the callers’ assumptions
How often will it run?Not answerableFrom call sites: loops, batch jobs, request paths
What does it hold while it runs?Only if the transaction is in the same hunkFrom callers’ transactions, locks, and pool usage
What is its worst-case latency?Only the timeout written in the diffTimeout × inherited retries, checked against caller deadlines
Is the fix complete?For the file shownAcross sibling copies
Does it bypass a rule?Only if the rule is in the diffBy locating where the rule is enforced

For the diff-level questions themselves (resource use, failure handling, downstream load), How to Review a Pull Request for Production Reliability Risks has the checklist. For weighing how far a change reaches once you know its callers, see How to Assess the Blast Radius of a Code Change.

What can a repository-wide review not see?

A repository-wide review cannot see runtime facts. It can tell you that a function is called in a loop; it cannot tell you the loop runs 50,000 times unless something in the code or its data says so. It is a large improvement over the diff and still a partial view.

GapWhy the repository cannot show itWhere to get it
Traffic and call frequencyVolumes are properties of production, not codeRequest metrics, traces, job run history
Data size and distributionTable sizes and skew change after the code is writtenDatabase statistics, query plans on production-like data
Dependency latencyThe tax service’s p99 is not in your repoClient-side latency metrics, SLOs of the dependency
Configuration outside the repoEnv vars, secrets, and infra may live in another systemDeployment manifests, infrastructure repositories, the owning team
Other services’ callersA public API’s consumers live in other repositoriesAPI gateway logs, consumer contracts, service catalog
Dynamically wired codeCallers registered by name or config are invisible to text searchHandler registries, DI configuration, scheduler config

That is why the strongest reviews combine all three views: the diff for intent, the repository for context, and production signals for scale. The last is covered in Production Reliability vs Code Review.

How Tomosu helps

Tomosu analyzes the standing codebase as well as each pull request, so the cross-file context in this post is available when a change is reviewed rather than reconstructed by hand. For a given PR, it surfaces:

These signals roll up into the Production Reliability Index, weighted by blast radius, so a change that puts a remote call inside the checkout transaction ranks above the same call in an internal script.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

What is a repository-wide code review?

A repository-wide review evaluates a change against the whole codebase rather than only the changed lines. It follows the callers and callees of changed code, the shared configuration it inherits, sibling copies of the same pattern, and invariants enforced in other files, and it can also scan code the pull request did not touch for standing risk.

Why does a PR diff miss production reliability risks?

Most reliability risk depends on context the diff does not contain: how often a function is called, whether a caller holds a transaction or lock, which retry and timeout policy a shared client applies, and what data volumes the code sees. A change can be correct line by line and still be unsafe because of code in another file.

Does a repository-wide review replace pull request review?

No. Diff review is still where reviewers judge intent, correctness, naming, and design. A repository-wide review adds the context the diff cannot show, so the reviewer can answer questions about callers, shared configuration, and blast radius with evidence instead of memory.

How do I find every caller of a changed function during code review?

Use your IDE’s find usages or call hierarchy, or run git grep -n -w with the function name across the repository. Check decorators, dependency injection, message handlers, and scheduled jobs separately, because text search and static call graphs both undercount callers that are wired by name or configuration.

What is an activated pre-existing finding?

It is a risk that already existed in code the pull request did not change, but that the change makes reachable, more frequent, or more severe. For example, a missing timeout in an old client becomes urgent when a PR starts calling that client on the checkout path. It belongs to the PR’s risk even though its lines are not in the diff.

Should pre-existing issues found by a repository-wide review block a pull request?

Usually not. Pre-existing issues the change does not touch should go to a tracked backlog with an owner, so they do not stall unrelated work. Findings the change introduces, and pre-existing findings the change activates, are the ones to weigh for this merge, using a severity policy agreed in advance.

Can AI code review tools see beyond the diff?

It depends on the tool. Some review only the diff and a few surrounding lines, while others index the repository and retrieve related files. Ask which files a finding was based on. A finding about a caller or a shared configuration should cite that file and line, not just the changed hunk.

What can a repository-wide review still not see?

It cannot see runtime facts: real traffic, data sizes, dependency latency, or configuration that lives outside the repository, such as environment variables set in another system. Those need production telemetry, deployment manifests, or a person who knows the system.


The diff is where a change is written. The repository is where it runs. Tomosu reads both before the change reaches production. Assess your repository →