A pull request diff is a list of changed lines. Whether those lines are safe usually depends on lines that did not change: the caller that holds a transaction, the shared client that retries, the job that calls the function fifty thousand times, the three other files that copy the same pattern. A reviewer who reads only the diff is judging a change without the code it runs inside.
A repository-wide review evaluates a change against the whole codebase, not just the changed hunks. It finds risks whose evidence lives in other files, which a PR diff cannot show even when every changed line is correct.
- Caller context: callers that loop, hold transactions or locks, or sit on hot paths.
- Inherited configuration: retries, timeouts, and pool sizes set elsewhere.
- Contract drift: changed return values, exceptions, or side effects that callers assume.
- Sibling copies: the same bug, fixed in one file and left in three others.
- Bypassed invariants and activated standing risk in code the PR newly reaches.
This is not an argument against reading diffs. Diff review is where intent, correctness, and design get judged, and nothing replaces it. The point is narrower: some of the most expensive production failures come from changes that look correct in the diff, because the reason they fail is in a file the diff never shows. This guide covers what those blind spots are, how to find them by hand, and how to classify what you find so it does not turn every pull request into an audit.
What does a PR diff actually show a reviewer?
A PR diff shows the lines that were added, removed, or modified, plus a few lines of surrounding context. On GitHub and GitLab the default view is the changed files, with the option to expand hidden lines inside those files. Files the pull request did not touch do not appear at all, however much they depend on the change.
That is a sensible design for reading a change. It is a poor design for judging a change’s effect on production, because production behavior is a property of the call path, not of the file. A function’s cost depends on who calls it, how often, and while holding what. A new network call’s worst case depends on the client it goes through. A validation rule’s strength depends on whether every write path goes through it.
A repository-wide review is a review that evaluates a change against the whole codebase: its callers and callees, the configuration it inherits, other code that follows the same pattern, and the invariants enforced elsewhere. It can also scan code the pull request did not touch, which matters when the change makes old code newly reachable.
How can a clean diff still break production?
Here is a complete pull request. A pricing function that used to read from an in-memory table now adds tax from a tax service. The author added a timeout, which is more than many reviewers would ask for. The tests pass.
def get_price(sku: str) -> Decimal:
base = _price_table[sku]
- return base
+ rate = http.get(f"{TAX_URL}/rates/{sku}", timeout=2).json()["rate"]
+ return base * (1 + Decimal(str(rate)))
Nothing in those lines is wrong. Now read two files that are not in the diff. The first is the shared HTTP session the new call goes through:
from requests import Session
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
http = Session()
# GETs retried on connect/read errors and 502/503/504
http.mount("https://", HTTPAdapter(max_retries=Retry(
total=2, backoff_factor=0.5, status_forcelist=[502, 503, 504])))
The second is the checkout path that calls get_price:
def place_order(session: Session, cart: Cart) -> Order:
with session.begin():
# connection held and stock rows locked from here...
stock = session.scalars(
select(Stock).where(Stock.sku.in_(cart.skus)).with_for_update()
).all()
reserve(stock, cart)
lines = [
OrderLine(sku=i.sku, qty=i.qty,
unit_price=get_price(i.sku)) # ...to here
for i in cart.items
]
order = Order(customer_id=cart.customer_id, lines=lines)
session.add(order)
return order
Put together, the change means every checkout holds a database connection and row locks on popular SKUs while it makes one tax call per cart item, each of which can be attempted three times with backoff. A healthy tax service adds tens of milliseconds per item. A slow one turns a 40 ms transaction into many seconds. Ten concurrent slow checkouts pin a ten-connection pool, and every other request that needs the database waits. Meanwhile the gateway gives up after 3 seconds, so the client may retry an order the server is still processing.
None of that is visible from pricing/service.py. The reviewer would have had to know that place_order calls get_price inside session.begin(), that the session retries, and that a nightly job calls the same function for every SKU in the catalog, turning one deploy into 50,000 or more tax calls per night. And the unit test mocks http.get, so it proves the arithmetic and nothing about latency. That is the pattern behind many PRs that pass CI but break production.
The fix is not in the diff’s file either. It is a design change at the caller: fetch prices in one batched call before the transaction opens, and give the batch job its own rate-limited path.
def place_order(session: Session, cart: Cart) -> Order:
# one batched call, before any connection or lock is held
prices = get_prices(cart.skus)
with session.begin():
stock = session.scalars(
select(Stock).where(Stock.sku.in_(cart.skus)).with_for_update()
).all()
reserve(stock, cart)
lines = [OrderLine(sku=i.sku, qty=i.qty, unit_price=prices[i.sku])
for i in cart.items]
order = Order(customer_id=cart.customer_id, lines=lines)
session.add(order)
return order
The diff tells you what changed. The repository tells you what the change will do.
What does a repository-wide review find that a PR diff misses?
A repository-wide review finds six kinds of risk that a PR diff cannot show, because their evidence lives in other files. Each has a different place to look.
1. Caller context: loops, transactions, locks, and hot paths
A changed function inherits its callers’ circumstances. A new database query is harmless in a settings page and an N+1 problem inside a loop. A new remote call is fine in a background worker and a pool-exhaustion risk inside a transaction. The question to answer is not “is this function correct?” but “who calls this, how often, and while holding what?” For the loop case in particular, see N+1 queries that pass tests but fail under production traffic.
2. Inherited configuration: retries, timeouts, pools, and flags
New code rarely configures its own client. It borrows a shared HTTP session, database pool, message producer, or cache client, and it inherits that object’s retry policy, timeouts, and limits. A diff that adds retries=3 looks modest until you find that the shared client already retries and the gateway retries too. That multiplication is the subject of How to Spot a Retry Storm Before Merge. The same applies to timeouts that no longer fit inside the caller’s deadline, and to feature flags whose default lives in a config file the PR did not touch.
3. Contract drift: return values, exceptions, and side effects
When a function starts returning None instead of raising, raising a new exception type, returning a lazily evaluated iterator instead of a list, or writing to a table it used to only read, every caller that relied on the old behavior is affected. The diff shows the new behavior. It does not show the caller that catches KeyError and now never sees it, or the caller that iterates the result twice.
4. Sibling copies: the same pattern in other files
Codebases repeat themselves. When a PR fixes a bug in one place, the same bug often exists in the files that were copied from it. When a PR introduces a pattern, a reviewer should know whether the repository already has a safer established version of it.
The webhook race in this figure is the check-then-insert pattern covered in Why “Check Then Insert” Creates Duplicate Records. It is a good example because the fix, a unique constraint plus an insert that handles conflicts, is short and local, which makes it easy to apply to one handler and forget the rest.
5. Bypassed invariants: rules enforced somewhere else
Many correctness rules are enforced in exactly one place: a service method that validates input, a repository class that adds the tenant filter, a wrapper that takes a lock or checks an idempotency key. A PR that adds a new write path, for example a bulk import endpoint that writes directly with the ORM, can skip that place without any line in the diff looking wrong. You only see the bypass by knowing where the rule is enforced and checking whether the new path goes through it.
6. Activated standing risk: old code the change makes reachable
Some risk has been sitting in the repository for years without causing an incident, because nothing important called it. A client with no timeout that was used by a monthly report. A query without an index on a table that used to be small. When a PR starts calling that code from a request path, the old risk becomes the new change’s risk, even though none of its lines are in the diff. This is the category most specific to repository-wide review, and the reason a standing scan of the codebase is worth having before any PR arrives.
| Blind spot | What the diff shows | Where the evidence lives |
|---|---|---|
| Caller context | A new query or remote call in a function | Every call site: loops, transactions, locks, request versus batch paths |
| Inherited configuration | A call through an existing client | Where the client, pool, retry policy, timeout, or flag default is defined, including deploy manifests |
| Contract drift | A changed return value, exception, or side effect | Callers that depend on the old behavior: except blocks, null checks, repeated iteration |
| Sibling copies | A fix or a new pattern in one file | Other files with the same code shape, often copied from each other |
| Bypassed invariants | A new write or read path | The single place that validates, filters by tenant, locks, or deduplicates |
| Activated standing risk | A new call to existing code | The callee’s own unguarded calls, missing indexes, and unbounded reads |
How should findings outside the diff be classified?
Findings outside the diff should be split into three groups: new, activated, and pre-existing. The split matters because a repository-wide review will find things, and if every finding is treated as the PR author’s problem, reviews stall and people stop trusting the tool.
- New: the PR introduced the risk. It is reviewed in this PR, like any diff comment, even when the evidence is in another file (the checkout example above).
- Activated: the risk existed before, but the PR makes it reachable, more frequent, or more severe. It belongs to this PR because the PR changes its consequences. The fix may be in the old code or in how the PR calls it.
- Pre-existing: the risk existed before and the PR does not change its exposure. It goes to a tracked backlog with an owner. Blocking an unrelated PR on it punishes the person who happened to open the file.
Which of the new and activated findings should actually stop a merge is a separate policy question, covered in Should a Reliability Finding Block a Merge? A Severity Framework.
A finding about a caller, a shared client, or a sibling copy is only checkable if it names the file and line it is based on. “This may be called in a loop” is a guess. “jobs/reprice_catalog.py:41 calls this once per SKU” is evidence. For how to check a finding before acting on it, see AI Code Review False Positives: How to Verify a Reliability Finding.
How do you review a pull request with repository context?
You review a pull request with repository context by starting from the changed symbols, not the changed files, and following each one outward. You do not need to do this for every PR. Do it for changes to shared code, hot paths, clients, data access, and anything that adds I/O.
- List the changed symbols. Write down the functions, classes, config keys, schemas, and message formats whose behavior changed, not just the files.
- Find every caller and consumer. Use find usages or a call hierarchy in your IDE, or
git grep. For each call site, note whether it runs in a loop, inside a transaction or lock, on a request path, or in a batch job. - Trace the configuration the change inherits. Find where the client, pool, retry policy, timeout, and feature flag default it uses are defined, including gateway and deployment files in the repository.
- Search for sibling copies. Search for the same code shape the PR fixes or introduces, and check whether the repository already has an established safer version.
- Check invariants enforced elsewhere. Identify where validation, tenant filtering, locking, and idempotency are enforced, and confirm the new path goes through them.
- Classify what you find. Mark each finding new, activated, or pre-existing, and keep pre-existing findings out of this PR’s review thread.
A few commands cover most of the manual work in a Git repository:
# function and class definitions touched by this branch
git diff --unified=0 origin/main...HEAD -- '*.py' \
| grep -E '^[+-][[:space:]]*(def|class) '
# every call site of a changed function, across the whole repository
git grep -n -w 'get_price' -- '*.py'
# where shared clients configure retries and timeouts
git grep -n -E 'Retry\(|max_retries|timeout=' -- '*.py' '*.yaml'
# sibling copies of the pattern being fixed
git grep -n -E '\.filter_by\(event_id=' -- 'webhooks/' 'jobs/'
Text search is a floor, not a ceiling. It misses callers wired by name: decorators, dependency injection, message handlers subscribed by topic, scheduled jobs configured in YAML, and reflection. Check those registries separately. Static call graphs have the same gap for dynamic dispatch. Developers ask about this regularly on Stack Overflow, and the honest answer is that no single tool finds every caller.
| Question | Diff-only review | Repository-wide review |
|---|---|---|
| Is the changed code correct? | Yes, this is what it is for | Same, plus the callers’ assumptions |
| How often will it run? | Not answerable | From call sites: loops, batch jobs, request paths |
| What does it hold while it runs? | Only if the transaction is in the same hunk | From callers’ transactions, locks, and pool usage |
| What is its worst-case latency? | Only the timeout written in the diff | Timeout × inherited retries, checked against caller deadlines |
| Is the fix complete? | For the file shown | Across sibling copies |
| Does it bypass a rule? | Only if the rule is in the diff | By locating where the rule is enforced |
For the diff-level questions themselves (resource use, failure handling, downstream load), How to Review a Pull Request for Production Reliability Risks has the checklist. For weighing how far a change reaches once you know its callers, see How to Assess the Blast Radius of a Code Change.
What can a repository-wide review not see?
A repository-wide review cannot see runtime facts. It can tell you that a function is called in a loop; it cannot tell you the loop runs 50,000 times unless something in the code or its data says so. It is a large improvement over the diff and still a partial view.
| Gap | Why the repository cannot show it | Where to get it |
|---|---|---|
| Traffic and call frequency | Volumes are properties of production, not code | Request metrics, traces, job run history |
| Data size and distribution | Table sizes and skew change after the code is written | Database statistics, query plans on production-like data |
| Dependency latency | The tax service’s p99 is not in your repo | Client-side latency metrics, SLOs of the dependency |
| Configuration outside the repo | Env vars, secrets, and infra may live in another system | Deployment manifests, infrastructure repositories, the owning team |
| Other services’ callers | A public API’s consumers live in other repositories | API gateway logs, consumer contracts, service catalog |
| Dynamically wired code | Callers registered by name or config are invisible to text search | Handler registries, DI configuration, scheduler config |
That is why the strongest reviews combine all three views: the diff for intent, the repository for context, and production signals for scale. The last is covered in Production Reliability vs Code Review.
How Tomosu helps
Tomosu analyzes the standing codebase as well as each pull request, so the cross-file context in this post is available when a change is reviewed rather than reconstructed by hand. For a given PR, it surfaces:
- Call paths through the change: which request routes, jobs, and consumers reach the changed code, and whether they reach it in a loop, inside a transaction, or on a hot path.
- Inherited reliability settings: the retry, timeout, and pool configuration the changed code runs under when those are defined in the repository, so a new call’s worst case can be judged.
- Findings with their file and line: cross-file findings cite the caller or configuration they are based on, so a reviewer can check them.
- Baseline versus change: because the repository is scanned before the PR, findings the PR introduces or activates can be told apart from standing issues that belong in a backlog.
These signals roll up into the Production Reliability Index, weighted by blast radius, so a change that puts a remote call inside the checkout transaction ranks above the same call in an internal script.
Scan your repository with Tomosu →
Key takeaways
- A PR diff shows changed lines. Production behavior is a property of the call path, which usually runs through files the diff does not show.
- A repository-wide review finds caller context, inherited configuration, contract drift, sibling copies, bypassed invariants, and activated standing risk.
- A correct diff can still be unsafe: a new network call becomes a pool and lock problem when a caller makes it inside a transaction.
- Start from changed symbols, find every caller, and trace the configuration the change inherits. Text search misses callers wired by name.
- Classify findings as new, activated, or pre-existing. Only the first two belong to the PR; the third goes to a tracked backlog.
- A cross-file finding is only useful if it cites the file and line it is based on.
- Repository context still cannot see traffic, data size, or dependency latency. Combine it with production signals.
Frequently asked questions
What is a repository-wide code review?
A repository-wide review evaluates a change against the whole codebase rather than only the changed lines. It follows the callers and callees of changed code, the shared configuration it inherits, sibling copies of the same pattern, and invariants enforced in other files, and it can also scan code the pull request did not touch for standing risk.
Why does a PR diff miss production reliability risks?
Most reliability risk depends on context the diff does not contain: how often a function is called, whether a caller holds a transaction or lock, which retry and timeout policy a shared client applies, and what data volumes the code sees. A change can be correct line by line and still be unsafe because of code in another file.
Does a repository-wide review replace pull request review?
No. Diff review is still where reviewers judge intent, correctness, naming, and design. A repository-wide review adds the context the diff cannot show, so the reviewer can answer questions about callers, shared configuration, and blast radius with evidence instead of memory.
How do I find every caller of a changed function during code review?
Use your IDE’s find usages or call hierarchy, or run git grep -n -w with the function name across the repository. Check decorators, dependency injection, message handlers, and scheduled jobs separately, because text search and static call graphs both undercount callers that are wired by name or configuration.
What is an activated pre-existing finding?
It is a risk that already existed in code the pull request did not change, but that the change makes reachable, more frequent, or more severe. For example, a missing timeout in an old client becomes urgent when a PR starts calling that client on the checkout path. It belongs to the PR’s risk even though its lines are not in the diff.
Should pre-existing issues found by a repository-wide review block a pull request?
Usually not. Pre-existing issues the change does not touch should go to a tracked backlog with an owner, so they do not stall unrelated work. Findings the change introduces, and pre-existing findings the change activates, are the ones to weigh for this merge, using a severity policy agreed in advance.
Can AI code review tools see beyond the diff?
It depends on the tool. Some review only the diff and a few surrounding lines, while others index the repository and retrieve related files. Ask which files a finding was based on. A finding about a caller or a shared configuration should cite that file and line, not just the changed hunk.
What can a repository-wide review still not see?
It cannot see runtime facts: real traffic, data sizes, dependency latency, or configuration that lives outside the repository, such as environment variables set in another system. Those need production telemetry, deployment manifests, or a person who knows the system.
The diff is where a change is written. The repository is where it runs. Tomosu reads both before the change reaches production. Assess your repository →