Company
About Tomosu
Platform
Platform & Agents Indexes How it works Solutions Pricing
Get Started
MCP Server VS Code — Plugin Installation Scan Your Repo — Guide Integrations · GitHub App Integrations · CodeRabbit MCP FAQ
Free Tools
Governance Impact
Resources
Blogs News Download / Free Trial Book a call →
Production Debugging · Concurrency

Race Conditions That Only Appear Under Production Load

Tomosu AI·14 min read·

A balance goes negative. An order ships twice. A counter is off by three after a busy afternoon. Nobody can reproduce it, the tests are green, and the logs show nothing that looks like an error. This is what a race condition looks like in production. It appears only when enough requests touch the same data at the same moment.

Quick answer

Race conditions appear only under production load because a race needs two operations on the same data inside the same short window between a read and a write. Tests rarely create that overlap. Production does, through concurrency, hot keys, retries, and replicas. To reproduce the race, aim concurrent workers at one key. To fix it, make the read and the write one atomic step.

This guide explains why concurrency bugs hide from tests, where production concurrency comes from even for a single user, how to turn an intermittent failure into a test that fails every time, and which fix fits which race. The duplicate-record variant, where two requests both check that a row does not exist and then both insert it, has its own guide: Why “Check Then Insert” Creates Duplicate Records.

What is a race condition in a web service?

A race condition is a bug where the result depends on the timing of operations that run at the same time. In a web service, it almost always has the same shape. Code reads shared state, makes a decision or computes a new value in application code, and writes the result back. Another request changes the state in between.

The most common version is the lost update. Two requests read the same balance, each subtracts its amount, and each writes its own result. The second write silently overwrites the first.

A LOST UPDATE: TWO CORRECT REQUESTS, ONE WRONG RESULT Request A Database Request B read 100 write 70 read 100 write 80 100 70 80 A’s window: 100 − 30 B’s window: 100 − 20, based on a stale read Expected 50 No error Each request is correct on its own. The bug exists only in the overlap of their read-to-write windows.
A lost update raises no exception. The only symptom is data that disagrees with the history of requests.

The code that produces this looks entirely reasonable in review, which is part of the problem:

accounts.js · Node.js + PostgreSQLlost update under concurrency
app.post('/accounts/:id/withdraw', async (req, res) => {
  const { id } = req.params;
  const amount = Number(req.body.amount);
  const { rows: [acct] } = await db.query(
    'SELECT balance FROM accounts WHERE id = $1', [id]);          // read
  if (acct.balance < amount) return res.status(409).send('insufficient funds');
  await db.query(
    'UPDATE accounts SET balance = $1 WHERE id = $2',
    [acct.balance - amount, id]);                                       // write, based on a stale read
  res.sendStatus(204);
});

Two concurrent withdrawals can both pass the balance check and then overwrite each other. The balance check does not protect anything, because it runs against a value that may already be out of date.

Why do race conditions only appear under production load?

Because a race needs an overlap, and overlaps scale with traffic on the same key and with the length of the window. A useful approximation: the number of other requests in flight on the same entity during one request’s read-to-write window is roughly the arrival rate on that entity multiplied by the window length. When that number is tiny, the race is theoretical. When it approaches 1, the race happens all the time.

OVERLAPS ≈ REQUESTS/S ON ONE KEY × READ-TO-WRITE WINDOW Local test, one key0.2 req/s × Window30 ms = Overlapping requests0.006: never seen Production hot key50 req/s × Window30 ms = Overlapping requests1.5: routine Same hot key50 req/s × Window, slow query or GC300 ms = Overlapping requests15: during incidents Both factors grow in production: traffic concentrates on hot keys, and slow moments stretch every window.
The same code has a race probability close to zero in tests and close to certainty on a hot key during a slow minute.

Every factor in that product is different in production:

That is also why these bugs look like they “come out of nowhere” after a harmless change. A new remote call inside the window, a slower query, or a marketing campaign that creates one very hot product can raise the overlap rate by orders of magnitude without touching the racy code.

Where does production concurrency come from, even for one user?

Teams often dismiss a race because “one user cannot do two things at once.” Production creates concurrent executions of the same operation for a single user all the time:

ONE CLICK, TWO CONCURRENT EXECUTIONS Client Server: req 1 Server: req 2 send timeout at 2 s, retry read order · slow query · ... still running · write read order · write overlap: same order, same moment 0 s2 s3.5 s
A client timeout shorter than the server’s worst-case latency turns every slow request into two concurrent ones.
SourceHow it creates concurrent executions
Client or gateway retriesThe client gives up and retries while the original is still running server-side. See how retries multiply load.
Double submitsDouble-clicks, mobile reconnects, and browser back-and-resubmit send the same action twice within milliseconds.
At-least-once deliveryQueues and webhooks redeliver, sometimes while the first delivery is still being processed. Covered in preventing duplicate webhook processing.
Scheduled jobs on every replicaA cron-style job defined in application code runs once per instance, so 6 replicas run it 6 times at the same minute.
Rolling deploymentsOld and new versions run together; a consumer rebalance can hand the same partition or lease to two processes briefly.
Background work and user actionsA nightly recalculation updates the same rows a user is editing.

What evidence does a race condition leave in production?

Races rarely throw. They leave data that is inconsistent with the requests that were made. The investigation is about finding two operations on the same entity whose time ranges overlap.

SymptomLikely race
Counter, balance, or stock slightly off; no errorsLost update from read-modify-write in application code
Negative stock or a limit exceededCheck-then-act: both requests passed the check before either wrote
Duplicate rows, emails, or chargesCheck-then-insert or a redelivered message processed concurrently
An “impossible” state, such as shipped but not paidOrdering race: an older update applied after a newer one
Problems cluster around slow periods or deploysLonger windows, retries, and overlapping versions raised the overlap rate

To confirm it, you need logs that let you reconstruct concurrent access to one entity:

The guide to correlating logs, traces, and a code change covers the tooling. For races, the key filter is the entity ID, not the request ID.

How do you reproduce a race condition that only happens in production?

You reproduce a race by recreating the overlap deliberately, not by waiting for it. The procedure:

  1. Name the shared state and the broken invariant. Which row, key, or object ended up wrong, and which rule did it break? “Balance never below zero”, “one shipment per order”.
  2. Find the read, the write, and the window between them. Measure how long the gap is in production, including queries and remote calls in between.
  3. List every source of concurrency on that key. Users, retries, replicas, consumers, jobs, deploys.
  4. Build a concurrent test on one key. Many workers, one entity, started together with a barrier, repeated many times.
  5. Widen the window if it does not reproduce. Inject a delay between read and write in the test build.
  6. Assert the invariant, fix, and rerun. The same test that failed must now pass consistently.
WithdrawRaceTest.java · JUnit 5, real database
@RepeatedTest(50)
void concurrentWithdrawalsNeverLoseUpdates() throws Exception {
    long id = accounts.create(1_000);                        // one hot key
    int workers = 32;
    CountDownLatch start = new CountDownLatch(1);
    ExecutorService pool = Executors.newFixedThreadPool(workers);
    try {
        List<Future<?>> done = new ArrayList<>();
        for (int i = 0; i < workers; i++) {
            done.add(pool.submit(() -> { start.await(); service.withdraw(id, 10); return null; }));
        }
        start.countDown();                                           // release all at once
        for (Future<?> f : done) f.get(10, TimeUnit.SECONDS);
    } finally {
        pool.shutdownNow();
    }
    assertEquals(1_000 - workers * 10, accounts.balance(id));      // the invariant, not "no errors"
}

Three details make the difference between a test that finds the race and one that passes by luck:

A flaky concurrency test is a finding

If a concurrent test passes most of the time and fails sometimes, do not mark it flaky and retry it. That intermittent failure is the race, caught cheaply. The cost of treating it as noise is an incident later.

Which fix fits which race condition?

Every correct fix does the same thing: it makes the read and the write indivisible, or it makes a conflicting write fail loudly instead of silently winning. Sleeps, retries, and “check again just before writing” only shrink the window. Shrinking the window lowers the frequency without removing the race.

PICK THE SMALLEST TOOL THAT MAKES IT ATOMIC Q1 · SHAPE OF THE CHANGE Can it be one conditional statement on one row? Q2 · KIND OF INVARIANT Is the invariant uniqueness (one row per key)? Q3 · CONTENTION Are conflicts on the same row rare? Atomic update UPDATE … WHERE; check rows Unique constraint INSERT … ON CONFLICT Optimistic locking Version column, retry YESYESYES NONONO Hot row: pessimistic lock or serialize per key SELECT … FOR UPDATE in a short transaction, or route all work for one key through one queue partition.
Start at the top. Most read-modify-write races in web services end at the first box.

Atomic conditional update

If the new value can be computed by the database from the old one, let the database do it in one statement. Put the condition in the WHERE clause and check how many rows changed:

accounts.jsread and write in one statement
const { rowCount } = await db.query(
  `UPDATE accounts
      SET balance = balance - $1
    WHERE id = $2 AND balance >= $1`,
  [amount, id]);
if (rowCount === 0) return res.status(409).send('insufficient funds');   // no stale read to act on
res.sendStatus(204);

This is safe at PostgreSQL’s default READ COMMITTED level. When two updates target the same row, the second waits for the first to commit. The PostgreSQL isolation documentation describes how the WHERE clause is then re-evaluated against the updated row. The second withdrawal sees the new balance, not the one it would have read earlier.

Optimistic locking with a version column

When the change needs application logic, such as a state machine transition or a document edit, read the row with its version and make the write conditional on the version being unchanged:

SQL · optimistic lockingconflict detected, not overwritten
UPDATE orders
   SET status = 'shipped', version = version + 1
 WHERE id = $1 AND version = $2;      -- $2 = the version you read
-- 0 rows updated: someone changed the order since you read it. Reload, re-decide, retry (bounded).

JPA and Hibernate do this for you with a @Version field. A conflicting write throws OptimisticLockException, which Spring translates to ObjectOptimisticLockingFailureException. Optimistic locking works well when conflicts are rare. On a hot row, most attempts conflict and the retries become their own load problem.

Pessimistic locking

SELECT ... FOR UPDATE locks the rows it reads until the transaction ends, so a second transaction waits instead of reading a value that is about to change. Keep the transaction short, and never make a remote call while holding the lock. The waiting requests each hold a database connection, which is how a lock turns into pool exhaustion. Lock rows in a consistent order to avoid deadlocks.

Stricter isolation

Under REPEATABLE READ or SERIALIZABLE, PostgreSQL aborts one of two conflicting transactions with could not serialize access due to concurrent update (or a serialization failure under SERIALIZABLE). That turns a silent lost update into an error, but only if the application retries the whole transaction. Raising the isolation level without adding that retry trades a data bug for an error rate.

A sleep, a retry, or a second check makes a race rarer. Only an atomic operation makes it impossible.

Races in memory: threads and async code

Not every race is in the database. Shared in-process state races too, and the tools differ by runtime.

Threads (Java, Go, and others)

count++ on a shared field is a read, an add, and a write, so concurrent threads lose increments. A plain HashMap modified from several threads can lose entries or corrupt its structure. Use the atomic operations the runtime provides: AtomicLong.incrementAndGet(), or ConcurrentHashMap.merge(key, 1L, Long::sum) instead of get followed by put. Go’s race detector (go test -race) and ThreadSanitizer find unsynchronized memory access at run time, and jcstress tests JVM concurrency behaviour. None of them can see a lost update between two requests that each talk to a database.

Async code (Node.js and friends)

“Node.js is single-threaded, so it has no races” is half true. Code between two awaits runs without interruption, but every await is a point where other requests run. A check before an await and an action after it is a check-then-act race, in one process with no threads:

tokens.jstwo refreshes in flight
let token = null;
async function getToken() {
  if (!token || token.expiresAt < Date.now()) {
    token = await fetchNewToken();     // 20 concurrent callers all see "expired" and all fetch
  }
  return token;
}

// Fixed: share the in-flight promise so only one refresh runs per process
let refreshing = null;
async function getToken() {
  if (token && token.expiresAt >= Date.now()) return token;
  refreshing ??= fetchNewToken().then(t => (token = t)).finally(() => { refreshing = null; });
  return refreshing;                   // callers await the same refresh
}

The in-flight promise fixes the race inside one process. With several processes or replicas, each still refreshes once. That is fine for a token, but not for anything that must happen exactly once, which needs the database or another shared atomic primitive. The same pattern at cache scale is covered in Cache Stampede: Why a Healthy Cache Can Overload Your Database.

In-process locks do not cross processes

A mutex, synchronized block, or per-key promise map protects state within one process. Once the service runs on more than one replica, it does nothing for shared data in a database or cache. Distributed locks exist but are subtle: a lock holder can pause past its lease and keep writing. Prefer constraints and conditional writes in the data store itself.

What should you look for in review to catch races before production?

Races are cheapest to catch when the racy shape is introduced, because the shape is visible in code even when the timing is not. In review, look for:

Racy shapes
  • A read, then a decision or computation in code, then a write of the same entity
  • An existence check followed by an insert
  • An await or remote call between a check and the action it guards
  • Shared mutable state in module scope, singletons, or static fields
Concurrency sources
  • Is the endpoint retried by clients or the gateway?
  • Is the job scheduled in app code on every replica?
  • Does the consumer tolerate redelivery during processing?
  • Which keys will be hot in production?
Proof of the fix
  • Atomic statement, version check, lock, or constraint named
  • Concurrent test on one key that asserts the invariant
  • Conflict path handled: 409, retry with a bound, or reload

These checks belong to the broader practice in How to Review a Pull Request for Production Reliability Risks. The difference with races is that the reviewer must reason about two executions at once, which is exactly what a single-diff view makes hard.

How Tomosu helps

A race is visible in the code long before it is visible in the data. Tomosu analyzes the repository and each pull request for the shapes that turn into production races, and for the context a diff does not show:

These findings feed the Production Reliability Index. Racy code gets discussed in review, with the concurrent test that would prove the fix, rather than reconstructed from inconsistent data weeks later.

Scan your repository with Tomosu →

Key takeaways

Frequently asked questions

Why do race conditions only appear under production load?

A race needs two operations on the same state inside the same short window. Tests send few requests, spread across unique keys, with fast dependencies, so overlaps almost never happen. Production has more concurrent requests, hot keys that receive a large share of traffic, retries, multiple replicas, and slow moments that widen the window.

How do I reproduce a race condition that only happens in production?

Aim many concurrent workers at the same entity, start them together with a barrier, and repeat the run many times. If it still does not fail, inject a delay between the read and the write in a test build to widen the window. Assert the invariant that broke in production, not just the absence of errors.

What is a lost update?

A lost update happens when two operations read the same value, each computes a new value from it, and both write. The second write overwrites the first, so one change disappears without any error. It is the most common race condition in web services that read a row, change it in application code, and write it back.

Can a race condition happen in single-threaded Node.js?

Yes. JavaScript code between two awaits runs without interruption, but every await lets other requests run. If a handler reads state, awaits something, and then writes based on what it read, another request can change the state in between. Multiple Node.js processes or replicas add true parallelism on top of that.

Should I use optimistic or pessimistic locking to fix a race?

Prefer a single atomic conditional update when the change fits in one statement. Use optimistic locking with a version column when conflicts are rare, and retry on conflict. Use pessimistic locking, such as SELECT FOR UPDATE, when the same row is contended often and retries would waste work, and keep that transaction short.

Do retries cause race conditions?

They often expose them. When a client times out and retries while the original request is still running on the server, one user action becomes two concurrent executions of the same operation. If that operation is a read-modify-write or a check-then-act, the retry can create a lost update or a duplicate.

Can the Go race detector or ThreadSanitizer find database race conditions?

No. They detect unsynchronized access to shared memory inside one process while the code runs. A lost update between two requests that each read and write a database row is not a memory data race, so it needs a concurrent integration test against a real database instead.


A race condition is two correct requests and one wrong result. Tomosu finds the read-to-write windows in your code before production traffic finds them for you. Assess your repository →