Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Deployment

A PR Passed CI but Broke Production: What Did the Tests Miss?

Tomosu AI·14 min read·

Every check was green. The reviewer approved. Twenty minutes after the deploy, error rates climbed and the team rolled back. Someone asks the obvious question in the incident channel: how did this get past CI? The answer is almost never “someone skipped the tests”. It is that the tests were answering a different question from the one production asked.

Quick answer

A PR that passes CI but breaks production usually fails on something CI does not model. CI checks the change against test data, test configuration, mocked dependencies, and one version at a time. Find which difference the failure depends on, then add the cheapest check at the earliest stage that can see it.

The instinct after an escape is to write a test for the exact bug. That closes one hole. This guide shows how to classify the escape so you close the class, and how to pick between a unit test, a contract test, a migration rehearsal, and a canary.

Why does a PR pass CI and still break production?

CI proves that the change behaves correctly for the inputs, configuration, and dependencies the test environment provides. It does not prove the change is safe in production, because production supplies different inputs, runs many requests at once, uses different configuration, talks to real dependencies, and deploys the change onto live data while the old version is still running.

A test escape is a defect that passed every automated check and was found in production. Every escape lives in one of the gaps below. Naming the gap is the first step, because it decides what kind of check could have caught it.

WHAT CI CHECKS VS WHAT PRODUCTION DOES DIMENSIONIN CI (GREEN)IN PRODUCTION Data shape Fixtures you wrote Nulls, legacy values, unicode, huge rows Data volume Tens of rows Millions of rows, different query plans Concurrency One request at a time Hundreds of overlapping requests Configuration Test profile, default flags Real env vars, secrets, flag states Dependencies Mocks that answer instantly Latency, timeouts, 429s, partial failures Deploy Fresh schema, one version Live-table migration, mixed versions A test escape lives in one of these rows. The row decides which check could have caught it.
CI and production run the same code under different conditions. Every escape depends on at least one row where they differ.

What kinds of failures do tests usually miss?

Grouping escapes by the gap they slipped through is more useful than grouping them by component. The same five classes show up across languages and stacks.

Escape classWhy CI misses itTypical exampleCheck that catches it
Missing inputFixtures contain only the cases the author imaginedA null middle_name, a 3 MB payload, an old enum value in a row from 2019Tests built from sampled, anonymized production shapes; property-based tests
ConfigurationCI uses a test profile and default flag valuesA flag on in production and off in CI; a missing env var that defaults to an empty stringStartup config validation; a pre-deploy config diff
Mock driftMocks return fast successes the author choseThe payment API times out or returns 429; the error path was never runFailure-injection tests; contract tests; integration tests against a real container
LoadTests run one request at a time on tiny tablesA race between two writers; a missing index that turns 20 ms into 4 s at scaleReview of shared-state and query changes; load tests; canary under real traffic
Deploy-timeCI builds a fresh schema and runs one versionA migration that locks a live table; new code writing a value old pods cannot readMigration rehearsal on production-sized data; expand and contract; mixed-version canary

Some escapes have nothing to do with the tests. If CI is flaky, the team learns to re-run until green, and a real failure can be dismissed as noise. That is a separate problem with its own cost, covered in What a Flaky Test Actually Costs You When AI Wrote the Code. Before classifying, check that the failing test (if there was one) was not simply retried into passing.

How do you find what the tests missed?

Start from the incident, not the test suite. You need the exact thing that failed in production and the exact conditions around it. If you do not yet know which PR caused the incident, narrow that down first; Which Commit Caused the Production Incident? walks through it.

  1. Capture the failing production input. The exact request, message, or row, plus the configuration, flag states, and service versions that were live at the time.
  2. Replay it in a unit or integration test. Run the exact input against the changed code. If it fails, the suite was missing that input or never asserted that outcome.
  3. Diff configuration and flags. Compare environment variables, secrets, feature flag states, and dependency versions between CI and production.
  4. Check the mocks against the real dependency. Compare what the mocks return with what the real service returns for errors, timeouts, rate limits, and slow responses.
  5. Check load and deploy conditions. Look for concurrency, data volume, migration locks, and old and new versions running side by side during the rollout.
  6. Add the cheapest check for that class. Put it at the earliest stage that can see the failure, and search for sibling code with the same gap.
WHICH GAP DID IT SLIP THROUGH? FOUR QUESTIONS, IN ORDER Q1 · REPLAY THE INPUT Does the exact production input fail in a unit test? Q2 · CONFIG AND FLAGS Does it fail only with production config or flags? Q3 · REAL DEPENDENCY Does it fail only against the real dependency? Q4 · LOAD Does it fail only under concurrency or volume? Missing input Add the case; find siblings Configuration Validate config at startup Mock drift Make mocks fail like the real one Load Races, locks, query plans YESYESYESYES NONONONO Deploy-time escape Migration locks, old and new versions side by side, or deploy order. Rehearse the rollout, not just the code.
The first “yes” names the gap. If every answer is “no”, the code was fine in isolation and the deploy itself broke it.

The last step matters most. An escape is rarely unique: the same missing timeout handling or the same unsafe migration pattern usually exists in several places. Fixing the one instance that fired leaves the others waiting for their turn.

Example: the mock that always succeeds

A checkout PR adds a payment call. The test mocks the payment client with a successful response, asserts a 201, and passes. In production the payment provider has a slow minute. The call times out, the handler has no path for that, the request fails with a 500, and the order row has already been written.

test_checkout.pyonly the happy path
def test_checkout_charges_card(client, mocker):
    mocker.patch("shop.checkout.payments.charge",
                 return_value=Charge(id="ch_1", status="succeeded"))
    resp = client.post("/checkout", json=CART)
    assert resp.status_code == 201

Coverage reports every line of the handler’s success path as covered. The except branch that does not exist cannot show up as uncovered. The fix is to make the mock fail in the ways the real dependency fails, and to assert what the system state should be afterwards, not only the status code:

test_checkout.pyreal failure modes, state asserted
@pytest.mark.parametrize("failure, expected_status", [
    (requests.Timeout(), 503),
    (requests.ConnectionError(), 503),
    (PaymentDeclined("card_declined"), 402),
])
def test_checkout_when_payment_fails(client, mocker, failure, expected_status):
    mocker.patch("shop.checkout.payments.charge", side_effect=failure)
    resp = client.post("/checkout", json=CART)
    assert resp.status_code == expected_status
    # never a paid order without a charge
    assert count_orders(status="paid") == 0

This is still a mock. It catches missing error handling, but not a change in the provider’s real response format. For that, use a contract test or an integration test against a container or sandbox of the real dependency. The point is to test the failure modes the dependency actually has, which you can read from its documentation, its status history, and your own traces.

A quick audit

Search your test suite for mocks of network clients and count how many ever use side_effect (or your framework’s equivalent) to raise a timeout or error. If the answer is close to zero, every error path that talks to a dependency is untested.

Example: the migration that was instant in CI

A PR adds an index to speed up a new query. In CI, the migration runs against an empty table in milliseconds. In production, the orders table has 40 million rows. A plain CREATE INDEX in PostgreSQL blocks inserts, updates, and deletes on the table until the build finishes, so checkout writes queue behind it for minutes.

0042_orders_customer_idx.sqlblocks writes on a large table
-- Instant on an empty CI database
CREATE INDEX idx_orders_customer ON orders (customer_id);
-- On 40M rows: writes to orders wait until the build completes
0042_orders_customer_idx.sqlonline, bounded wait
-- Must run outside a transaction block (turn off the tool's wrapper)
SET lock_timeout = '5s';
CREATE INDEX CONCURRENTLY IF NOT EXISTS idx_orders_customer
    ON orders (customer_id);
-- If it fails, it leaves an INVALID index: drop it and retry

The PostgreSQL CREATE INDEX documentation describes both the write lock of a normal build and the trade-offs of CONCURRENTLY. Other operations that behave very differently on a large table include adding a NOT NULL constraint to an existing column (a full table scan under a strong lock) and changing a column’s type (often a full table rewrite). None of these show up in CI, because CI has no data.

The check that catches this class is a migration rehearsal: run each migration against a restored, production-sized copy, record its duration and the locks it takes, and fail the pipeline when it exceeds a budget. Reviewing the PR for the missing index in the first place is a separate question, covered in Missing Database Indexes in a PR: What Reviewers Should Check.

Example: old and new versions running at the same time

A PR adds a new order status, PARTIALLY_REFUNDED. Every test passes, because every test runs the new code against itself. During the rolling deploy, new pods start writing the new status to the database and to a queue. Old pods, still serving traffic, read those rows and messages and fail to deserialize them:

Old pod logduring the rollout
com.fasterxml.jackson.databind.exc.InvalidFormatException:
  Cannot deserialize value of type `com.example.OrderStatus`
  from String "PARTIALLY_REFUNDED":
  not one of the values accepted for Enum class: [NEW, PAID, SHIPPED, REFUNDED]

The same thing happens in reverse after a rollback: the old version comes back and finds the new value already stored. No single-version test can see this, because the failure needs two versions and shared data.

ADDING A VALUE THAT OLD CODE MUST READ ONE-STEP RELEASE New pods (v2) Write the new status to the table and the queue Old pods (v1), still live Read the new value and throw on deserialization Rollback to v1 Stored rows and messages still hold the new value TWO-STEP RELEASE (READERS FIRST) Release A: readers accept it Add the value; map unknown values to a default. Nothing writes the new value yet. Safe to roll back: no new data exists. Release B: writers emit it Deploy only after A runs everywhere. Every running reader understands it. Safe to roll back to A. Tests run one version against itself. The rollout runs two versions against shared data. Order the change so every pair of adjacent versions can read what the other writes.
The two-step release costs one extra deploy. It removes the window where old code meets data it has never seen.

In a Jackson-based Java service, the reader side of release A can be as small as a default constant plus one setting:

Release A · OrderStatus.java + application.propertiestolerant reader
public enum OrderStatus {
    NEW, PAID, SHIPPED, REFUNDED, PARTIALLY_REFUNDED,
    @JsonEnumDefaultValue UNKNOWN   // any value this version does not know
}

# application.properties
spring.jackson.deserialization.read-unknown-enum-values-using-default-value=true

The same readers-first ordering applies to database columns (expand, migrate, contract), message schemas, API fields, and cache keys. Cache keys are a common case of their own, described in Stale Data After a Release: How to Debug Cache Invalidation.

Why doesn’t more test coverage fix this?

Line coverage measures which code ran during tests, not which inputs were tried or which outcomes were checked. All three examples above had high coverage. The checkout handler’s lines all ran. The migration ran. The enum was deserialized in a test. Coverage cannot report a missing branch, a condition that only exists at scale, or a second version of the service that was never in the test.

Green CI means the change is correct for the inputs you thought of. Production is made of the inputs you didn’t.

Chasing a coverage number after an escape tends to add tests of the kind you already have. The more useful questions are: which of the five gaps does our pipeline not model at all, and what is the cheapest stage that could?

Which release checks catch what CI misses?

Each stage of a delivery pipeline can see some gaps and is blind to others. Unit tests are cheap and fast but cannot see load or deploy order. A canary sees real traffic, but only after the change is live for some users. Put each check at the earliest stage that can actually observe the failure.

WHICH STAGE CAN SEE WHICH ESCAPE (TYPICAL) PR REVIEWUNIT TESTSINTEGRATIONPRE-DEPLOYCANARY Missing inputConfigurationMock driftConcurrencyData volumeMigration locksMixed versions Main place to catch it Can catch it sometimes Empty: that stage cannot see this gap
Unit tests own one row. The rest need review, a realistic environment, or controlled exposure to real traffic.
CheckWhat it addsEscape classes it closes
Failure-injection unit testsMocks that raise timeouts, errors, and rate limitsMock drift (error handling), missing input
Contract or container testsReal dependency behavior and response formatsMock drift
Startup config validationFail fast on missing or invalid settingsConfiguration
Migration rehearsalDuration and locks on production-sized dataData volume, migration locks
Mixed-version test or canaryOld and new versions against shared dataMixed versions, deploy order
Canary with automatic rollbackReal traffic mix and concurrency on a small sliceLoad, configuration, anything left
Reliability-focused PR reviewLooks at shared state, dependencies, and deploy order beyond the diffConcurrency, migrations, mixed versions

Concurrency escapes deserve a note. They are the hardest to reproduce after the fact and the easiest to spot in code: a read-then-write on shared state, a missing unique constraint, a counter updated without a lock. Race Conditions That Only Appear Under Production Load covers how to reproduce them; review is usually where they can be caught cheaply. The production reliability PR review checklist lists what to look for, and How to Identify High Risk Pull Requests helps decide which PRs deserve that extra attention.

What should a reviewer ask when CI is green?

Inputs and dependencies
  • What real inputs could this code see that no fixture has?
  • What happens when each new dependency call times out or errors?
  • Do the mocks ever fail?
Load and data
  • Is there shared state written by more than one request?
  • Do new queries have an index at production size?
  • How long will each migration take on the real table?
Deploy and rollback
  • Can the old version read everything the new one writes?
  • Does this depend on a flag or config value set elsewhere?
  • Is it safe to roll back after it has run for an hour?

The last question connects the escape to the incident response. If the change is not safe to roll back, the team needs to know before the deploy, not during it. How to Decide Whether to Roll Back a Release covers what makes a rollback unsafe.

How Tomosu helps

Tomosu does not replace your tests. It looks at a change the way production will experience it, which is where CI is blind. For each pull request, analyzed against the rest of the repository, it surfaces:

These signals feed the Production Reliability Index, so a PR that is green in CI but risky in production is visible as such before merge.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

Why did my PR pass CI but fail in production?

Because CI tests the change against the inputs, data, configuration, and dependencies of the test environment. Production differs in data shape and volume, concurrency, configuration and feature flags, real dependency behavior, and deploy conditions such as migrations on live tables and old and new versions running together. The failure lives in one of those differences.

What is a test escape?

A test escape is a defect that passed every automated check and was found in production. Classifying each escape (missing input, configuration, mock drift, load, or deploy-time) shows which kind of check your pipeline is missing, which is more useful than adding one more test for the exact bug.

Does higher test coverage prevent production failures?

Not by itself. Line coverage shows which code ran during tests, not which inputs were tried or which outcomes were asserted. A line can be covered by a test that never exercises the failing input, never asserts the result, or runs against a mock that cannot fail the way production does.

Why do mocks let production bugs through?

A mock returns whatever the test author expected, usually a fast success. The real dependency also times out, rate limits, returns partial errors, and changes its responses over time. If no test makes the mock fail the way the real service fails, the error-handling code is never exercised before production.

How do I test a database migration before production?

Run it against a copy of production-sized data, not an empty CI database, and measure how long it takes and which locks it holds. Set a lock timeout, use online operations such as CREATE INDEX CONCURRENTLY in PostgreSQL, and check that both the old and new application versions work with the new schema.

How do I catch bugs that only happen during a rolling deploy?

Treat the old and new versions as two clients of the same data and messages. Make readers tolerate new values before any writer emits them, use expand and contract for schema changes, and run a test or canary stage where both versions handle traffic together.

Should every production incident get a new test?

Every incident should get a check, but not always a unit test. If the escape came from configuration, data volume, or deploy order, the right check is a config validation, a migration rehearsal on real-sized data, or a canary with automatic rollback. Add the check at the earliest stage that can actually see the failure.


Green CI means the change works for the inputs someone thought of. Tomosu looks at what the change touches in production terms: shared data, dependencies, migrations, and the code around the diff. Assess your repository →