Every check was green. The reviewer approved. Twenty minutes after the deploy, error rates climbed and the team rolled back. Someone asks the obvious question in the incident channel: how did this get past CI? The answer is almost never “someone skipped the tests”. It is that the tests were answering a different question from the one production asked.
A PR that passes CI but breaks production usually fails on something CI does not model. CI checks the change against test data, test configuration, mocked dependencies, and one version at a time. Find which difference the failure depends on, then add the cheapest check at the earliest stage that can see it.
- Missing input: real data has nulls, legacy values, or sizes no fixture had.
- Configuration: production env vars, secrets, or flag states differ.
- Mock drift: the mock never fails the way the real dependency does.
- Load: concurrency and data volume expose races, locks, and slow plans.
- Deploy-time: migrations on live tables and mixed versions during rollout.
The instinct after an escape is to write a test for the exact bug. That closes one hole. This guide shows how to classify the escape so you close the class, and how to pick between a unit test, a contract test, a migration rehearsal, and a canary.
Why does a PR pass CI and still break production?
CI proves that the change behaves correctly for the inputs, configuration, and dependencies the test environment provides. It does not prove the change is safe in production, because production supplies different inputs, runs many requests at once, uses different configuration, talks to real dependencies, and deploys the change onto live data while the old version is still running.
A test escape is a defect that passed every automated check and was found in production. Every escape lives in one of the gaps below. Naming the gap is the first step, because it decides what kind of check could have caught it.
What kinds of failures do tests usually miss?
Grouping escapes by the gap they slipped through is more useful than grouping them by component. The same five classes show up across languages and stacks.
| Escape class | Why CI misses it | Typical example | Check that catches it |
|---|---|---|---|
| Missing input | Fixtures contain only the cases the author imagined | A null middle_name, a 3 MB payload, an old enum value in a row from 2019 | Tests built from sampled, anonymized production shapes; property-based tests |
| Configuration | CI uses a test profile and default flag values | A flag on in production and off in CI; a missing env var that defaults to an empty string | Startup config validation; a pre-deploy config diff |
| Mock drift | Mocks return fast successes the author chose | The payment API times out or returns 429; the error path was never run | Failure-injection tests; contract tests; integration tests against a real container |
| Load | Tests run one request at a time on tiny tables | A race between two writers; a missing index that turns 20 ms into 4 s at scale | Review of shared-state and query changes; load tests; canary under real traffic |
| Deploy-time | CI builds a fresh schema and runs one version | A migration that locks a live table; new code writing a value old pods cannot read | Migration rehearsal on production-sized data; expand and contract; mixed-version canary |
Some escapes have nothing to do with the tests. If CI is flaky, the team learns to re-run until green, and a real failure can be dismissed as noise. That is a separate problem with its own cost, covered in What a Flaky Test Actually Costs You When AI Wrote the Code. Before classifying, check that the failing test (if there was one) was not simply retried into passing.
How do you find what the tests missed?
Start from the incident, not the test suite. You need the exact thing that failed in production and the exact conditions around it. If you do not yet know which PR caused the incident, narrow that down first; Which Commit Caused the Production Incident? walks through it.
- Capture the failing production input. The exact request, message, or row, plus the configuration, flag states, and service versions that were live at the time.
- Replay it in a unit or integration test. Run the exact input against the changed code. If it fails, the suite was missing that input or never asserted that outcome.
- Diff configuration and flags. Compare environment variables, secrets, feature flag states, and dependency versions between CI and production.
- Check the mocks against the real dependency. Compare what the mocks return with what the real service returns for errors, timeouts, rate limits, and slow responses.
- Check load and deploy conditions. Look for concurrency, data volume, migration locks, and old and new versions running side by side during the rollout.
- Add the cheapest check for that class. Put it at the earliest stage that can see the failure, and search for sibling code with the same gap.
The last step matters most. An escape is rarely unique: the same missing timeout handling or the same unsafe migration pattern usually exists in several places. Fixing the one instance that fired leaves the others waiting for their turn.
Example: the mock that always succeeds
A checkout PR adds a payment call. The test mocks the payment client with a successful response, asserts a 201, and passes. In production the payment provider has a slow minute. The call times out, the handler has no path for that, the request fails with a 500, and the order row has already been written.
def test_checkout_charges_card(client, mocker):
mocker.patch("shop.checkout.payments.charge",
return_value=Charge(id="ch_1", status="succeeded"))
resp = client.post("/checkout", json=CART)
assert resp.status_code == 201
Coverage reports every line of the handler’s success path as covered. The except branch that does not exist cannot show up as uncovered. The fix is to make the mock fail in the ways the real dependency fails, and to assert what the system state should be afterwards, not only the status code:
@pytest.mark.parametrize("failure, expected_status", [
(requests.Timeout(), 503),
(requests.ConnectionError(), 503),
(PaymentDeclined("card_declined"), 402),
])
def test_checkout_when_payment_fails(client, mocker, failure, expected_status):
mocker.patch("shop.checkout.payments.charge", side_effect=failure)
resp = client.post("/checkout", json=CART)
assert resp.status_code == expected_status
# never a paid order without a charge
assert count_orders(status="paid") == 0
This is still a mock. It catches missing error handling, but not a change in the provider’s real response format. For that, use a contract test or an integration test against a container or sandbox of the real dependency. The point is to test the failure modes the dependency actually has, which you can read from its documentation, its status history, and your own traces.
Search your test suite for mocks of network clients and count how many ever use side_effect (or your framework’s equivalent) to raise a timeout or error. If the answer is close to zero, every error path that talks to a dependency is untested.
Example: the migration that was instant in CI
A PR adds an index to speed up a new query. In CI, the migration runs against an empty table in milliseconds. In production, the orders table has 40 million rows. A plain CREATE INDEX in PostgreSQL blocks inserts, updates, and deletes on the table until the build finishes, so checkout writes queue behind it for minutes.
-- Instant on an empty CI database
CREATE INDEX idx_orders_customer ON orders (customer_id);
-- On 40M rows: writes to orders wait until the build completes
-- Must run outside a transaction block (turn off the tool's wrapper)
SET lock_timeout = '5s';
CREATE INDEX CONCURRENTLY IF NOT EXISTS idx_orders_customer
ON orders (customer_id);
-- If it fails, it leaves an INVALID index: drop it and retry
The PostgreSQL CREATE INDEX documentation describes both the write lock of a normal build and the trade-offs of CONCURRENTLY. Other operations that behave very differently on a large table include adding a NOT NULL constraint to an existing column (a full table scan under a strong lock) and changing a column’s type (often a full table rewrite). None of these show up in CI, because CI has no data.
The check that catches this class is a migration rehearsal: run each migration against a restored, production-sized copy, record its duration and the locks it takes, and fail the pipeline when it exceeds a budget. Reviewing the PR for the missing index in the first place is a separate question, covered in Missing Database Indexes in a PR: What Reviewers Should Check.
Example: old and new versions running at the same time
A PR adds a new order status, PARTIALLY_REFUNDED. Every test passes, because every test runs the new code against itself. During the rolling deploy, new pods start writing the new status to the database and to a queue. Old pods, still serving traffic, read those rows and messages and fail to deserialize them:
com.fasterxml.jackson.databind.exc.InvalidFormatException:
Cannot deserialize value of type `com.example.OrderStatus`
from String "PARTIALLY_REFUNDED":
not one of the values accepted for Enum class: [NEW, PAID, SHIPPED, REFUNDED]
The same thing happens in reverse after a rollback: the old version comes back and finds the new value already stored. No single-version test can see this, because the failure needs two versions and shared data.
In a Jackson-based Java service, the reader side of release A can be as small as a default constant plus one setting:
public enum OrderStatus {
NEW, PAID, SHIPPED, REFUNDED, PARTIALLY_REFUNDED,
@JsonEnumDefaultValue UNKNOWN // any value this version does not know
}
# application.properties
spring.jackson.deserialization.read-unknown-enum-values-using-default-value=true
The same readers-first ordering applies to database columns (expand, migrate, contract), message schemas, API fields, and cache keys. Cache keys are a common case of their own, described in Stale Data After a Release: How to Debug Cache Invalidation.
Why doesn’t more test coverage fix this?
Line coverage measures which code ran during tests, not which inputs were tried or which outcomes were checked. All three examples above had high coverage. The checkout handler’s lines all ran. The migration ran. The enum was deserialized in a test. Coverage cannot report a missing branch, a condition that only exists at scale, or a second version of the service that was never in the test.
Green CI means the change is correct for the inputs you thought of. Production is made of the inputs you didn’t.
Chasing a coverage number after an escape tends to add tests of the kind you already have. The more useful questions are: which of the five gaps does our pipeline not model at all, and what is the cheapest stage that could?
Which release checks catch what CI misses?
Each stage of a delivery pipeline can see some gaps and is blind to others. Unit tests are cheap and fast but cannot see load or deploy order. A canary sees real traffic, but only after the change is live for some users. Put each check at the earliest stage that can actually observe the failure.
| Check | What it adds | Escape classes it closes |
|---|---|---|
| Failure-injection unit tests | Mocks that raise timeouts, errors, and rate limits | Mock drift (error handling), missing input |
| Contract or container tests | Real dependency behavior and response formats | Mock drift |
| Startup config validation | Fail fast on missing or invalid settings | Configuration |
| Migration rehearsal | Duration and locks on production-sized data | Data volume, migration locks |
| Mixed-version test or canary | Old and new versions against shared data | Mixed versions, deploy order |
| Canary with automatic rollback | Real traffic mix and concurrency on a small slice | Load, configuration, anything left |
| Reliability-focused PR review | Looks at shared state, dependencies, and deploy order beyond the diff | Concurrency, migrations, mixed versions |
Concurrency escapes deserve a note. They are the hardest to reproduce after the fact and the easiest to spot in code: a read-then-write on shared state, a missing unique constraint, a counter updated without a lock. Race Conditions That Only Appear Under Production Load covers how to reproduce them; review is usually where they can be caught cheaply. The production reliability PR review checklist lists what to look for, and How to Identify High Risk Pull Requests helps decide which PRs deserve that extra attention.
What should a reviewer ask when CI is green?
- What real inputs could this code see that no fixture has?
- What happens when each new dependency call times out or errors?
- Do the mocks ever fail?
- Is there shared state written by more than one request?
- Do new queries have an index at production size?
- How long will each migration take on the real table?
- Can the old version read everything the new one writes?
- Does this depend on a flag or config value set elsewhere?
- Is it safe to roll back after it has run for an hour?
The last question connects the escape to the incident response. If the change is not safe to roll back, the team needs to know before the deploy, not during it. How to Decide Whether to Roll Back a Release covers what makes a rollback unsafe.
How Tomosu helps
Tomosu does not replace your tests. It looks at a change the way production will experience it, which is where CI is blind. For each pull request, analyzed against the rest of the repository, it surfaces:
- New dependency calls without failure handling: remote calls added without a timeout, error path, or retry policy, which is where mock drift hides.
- Shared-state and concurrency risks: read-then-write patterns, missing constraints, and state touched by more than one request path.
- Migration and compatibility risks: schema changes that lock large tables, and new persisted values or formats that older code reads.
- Blast radius: which endpoints, jobs, and consumers depend on the changed code, so a green PR on a critical path gets the review it needs.
These signals feed the Production Reliability Index, so a PR that is green in CI but risky in production is visible as such before merge.
Scan your repository with Tomosu →
Key takeaways
- CI proves a change works under test conditions. Production breaks on the conditions CI does not model.
- Classify every escape: missing input, configuration, mock drift, load, or deploy-time. The class tells you which check is missing.
- Mocks that only succeed leave every error path untested. Make them fail the way the real dependency fails, and assert state.
- Migrations and mixed versions cannot fail in CI because CI has no data and one version. Rehearse migrations on real-sized data and release readers first.
- Coverage measures lines run, not inputs tried. It cannot show a branch that does not exist.
- Add each new check at the earliest stage that can see the failure, and fix the sibling code with the same gap.
Frequently asked questions
Why did my PR pass CI but fail in production?
Because CI tests the change against the inputs, data, configuration, and dependencies of the test environment. Production differs in data shape and volume, concurrency, configuration and feature flags, real dependency behavior, and deploy conditions such as migrations on live tables and old and new versions running together. The failure lives in one of those differences.
What is a test escape?
A test escape is a defect that passed every automated check and was found in production. Classifying each escape (missing input, configuration, mock drift, load, or deploy-time) shows which kind of check your pipeline is missing, which is more useful than adding one more test for the exact bug.
Does higher test coverage prevent production failures?
Not by itself. Line coverage shows which code ran during tests, not which inputs were tried or which outcomes were asserted. A line can be covered by a test that never exercises the failing input, never asserts the result, or runs against a mock that cannot fail the way production does.
Why do mocks let production bugs through?
A mock returns whatever the test author expected, usually a fast success. The real dependency also times out, rate limits, returns partial errors, and changes its responses over time. If no test makes the mock fail the way the real service fails, the error-handling code is never exercised before production.
How do I test a database migration before production?
Run it against a copy of production-sized data, not an empty CI database, and measure how long it takes and which locks it holds. Set a lock timeout, use online operations such as CREATE INDEX CONCURRENTLY in PostgreSQL, and check that both the old and new application versions work with the new schema.
How do I catch bugs that only happen during a rolling deploy?
Treat the old and new versions as two clients of the same data and messages. Make readers tolerate new values before any writer emits them, use expand and contract for schema changes, and run a test or canary stage where both versions handle traffic together.
Should every production incident get a new test?
Every incident should get a check, but not always a unit test. If the escape came from configuration, data volume, or deploy order, the right check is a config validation, a migration rehearsal on real-sized data, or a canary with automatic rollback. Add the check at the earliest stage that can actually see the failure.
Green CI means the change works for the inputs someone thought of. Tomosu looks at what the change touches in production terms: shared data, dependencies, migrations, and the code around the diff. Assess your repository →