A balance goes negative. An order ships twice. A counter is off by three after a busy afternoon. Nobody can reproduce it, the tests are green, and the logs show nothing that looks like an error. This is what a race condition looks like in production. It appears only when enough requests touch the same data at the same moment.
Race conditions appear only under production load because a race needs two operations on the same data inside the same short window between a read and a write. Tests rarely create that overlap. Production does, through concurrency, hot keys, retries, and replicas. To reproduce the race, aim concurrent workers at one key. To fix it, make the read and the write one atomic step.
- Why load matters: expected overlaps ≈ requests per second on one key × the read-to-write window.
- Reproduce: same entity, barrier start, many workers, many runs, and a widened window if needed.
- Fix: atomic conditional update, version check, row lock, or unique constraint. Not a sleep, not a retry.
This guide explains why concurrency bugs hide from tests, where production concurrency comes from even for a single user, how to turn an intermittent failure into a test that fails every time, and which fix fits which race. The duplicate-record variant, where two requests both check that a row does not exist and then both insert it, has its own guide: Why “Check Then Insert” Creates Duplicate Records.
What is a race condition in a web service?
A race condition is a bug where the result depends on the timing of operations that run at the same time. In a web service, it almost always has the same shape. Code reads shared state, makes a decision or computes a new value in application code, and writes the result back. Another request changes the state in between.
The most common version is the lost update. Two requests read the same balance, each subtracts its amount, and each writes its own result. The second write silently overwrites the first.
The code that produces this looks entirely reasonable in review, which is part of the problem:
app.post('/accounts/:id/withdraw', async (req, res) => {
const { id } = req.params;
const amount = Number(req.body.amount);
const { rows: [acct] } = await db.query(
'SELECT balance FROM accounts WHERE id = $1', [id]); // read
if (acct.balance < amount) return res.status(409).send('insufficient funds');
await db.query(
'UPDATE accounts SET balance = $1 WHERE id = $2',
[acct.balance - amount, id]); // write, based on a stale read
res.sendStatus(204);
});
Two concurrent withdrawals can both pass the balance check and then overwrite each other. The balance check does not protect anything, because it runs against a value that may already be out of date.
Why do race conditions only appear under production load?
Because a race needs an overlap, and overlaps scale with traffic on the same key and with the length of the window. A useful approximation: the number of other requests in flight on the same entity during one request’s read-to-write window is roughly the arrival rate on that entity multiplied by the window length. When that number is tiny, the race is theoretical. When it approaches 1, the race happens all the time.
Every factor in that product is different in production:
- Hot keys. Tests and load tests usually generate unique random IDs, which spreads traffic thinly. Real traffic concentrates on a few popular products, a large tenant, a shared counter, or one busy account. A system at 2,000 requests per second may send 50 of them to one row.
- Longer windows. The gap between read and write includes every query, remote call, and computation in between. A 30 ms window in normal conditions becomes 300 ms during a GC pause, lock wait, or slow dependency, which is exactly when traffic backs up.
- Real parallelism. A local run is often one process with a small pool. Production runs many replicas, each with many threads or connections.
- Correlated arrivals. Requests are not spread evenly. Double-clicks, client retries, batch jobs, and webhook bursts arrive together, which is the worst case for a race.
That is also why these bugs look like they “come out of nowhere” after a harmless change. A new remote call inside the window, a slower query, or a marketing campaign that creates one very hot product can raise the overlap rate by orders of magnitude without touching the racy code.
Where does production concurrency come from, even for one user?
Teams often dismiss a race because “one user cannot do two things at once.” Production creates concurrent executions of the same operation for a single user all the time:
| Source | How it creates concurrent executions |
|---|---|
| Client or gateway retries | The client gives up and retries while the original is still running server-side. See how retries multiply load. |
| Double submits | Double-clicks, mobile reconnects, and browser back-and-resubmit send the same action twice within milliseconds. |
| At-least-once delivery | Queues and webhooks redeliver, sometimes while the first delivery is still being processed. Covered in preventing duplicate webhook processing. |
| Scheduled jobs on every replica | A cron-style job defined in application code runs once per instance, so 6 replicas run it 6 times at the same minute. |
| Rolling deployments | Old and new versions run together; a consumer rebalance can hand the same partition or lease to two processes briefly. |
| Background work and user actions | A nightly recalculation updates the same rows a user is editing. |
What evidence does a race condition leave in production?
Races rarely throw. They leave data that is inconsistent with the requests that were made. The investigation is about finding two operations on the same entity whose time ranges overlap.
| Symptom | Likely race |
|---|---|
| Counter, balance, or stock slightly off; no errors | Lost update from read-modify-write in application code |
| Negative stock or a limit exceeded | Check-then-act: both requests passed the check before either wrote |
| Duplicate rows, emails, or charges | Check-then-insert or a redelivered message processed concurrently |
| An “impossible” state, such as shipped but not paid | Ordering race: an older update applied after a newer one |
| Problems cluster around slow periods or deploys | Longer windows, retries, and overlapping versions raised the overlap rate |
To confirm it, you need logs that let you reconstruct concurrent access to one entity:
- Log the entity ID and the value read at the read and at the write, with a request ID. Two requests that read the same version and both wrote are the proof.
- Query by entity, sorted by time. Pull every log line and trace span for the affected order or account around the incident. Overlapping start and end times on the same entity are the signal.
- Check the database side. In PostgreSQL,
log_lock_waitslogs statements that wait longer thandeadlock_timeoutfor a lock, which shows where requests contend on the same rows. - Correlate with retries and latency. A spike of client timeouts just before the inconsistent writes points at retry-induced concurrency.
The guide to correlating logs, traces, and a code change covers the tooling. For races, the key filter is the entity ID, not the request ID.
How do you reproduce a race condition that only happens in production?
You reproduce a race by recreating the overlap deliberately, not by waiting for it. The procedure:
- Name the shared state and the broken invariant. Which row, key, or object ended up wrong, and which rule did it break? “Balance never below zero”, “one shipment per order”.
- Find the read, the write, and the window between them. Measure how long the gap is in production, including queries and remote calls in between.
- List every source of concurrency on that key. Users, retries, replicas, consumers, jobs, deploys.
- Build a concurrent test on one key. Many workers, one entity, started together with a barrier, repeated many times.
- Widen the window if it does not reproduce. Inject a delay between read and write in the test build.
- Assert the invariant, fix, and rerun. The same test that failed must now pass consistently.
@RepeatedTest(50)
void concurrentWithdrawalsNeverLoseUpdates() throws Exception {
long id = accounts.create(1_000); // one hot key
int workers = 32;
CountDownLatch start = new CountDownLatch(1);
ExecutorService pool = Executors.newFixedThreadPool(workers);
try {
List<Future<?>> done = new ArrayList<>();
for (int i = 0; i < workers; i++) {
done.add(pool.submit(() -> { start.await(); service.withdraw(id, 10); return null; }));
}
start.countDown(); // release all at once
for (Future<?> f : done) f.get(10, TimeUnit.SECONDS);
} finally {
pool.shutdownNow();
}
assertEquals(1_000 - workers * 10, accounts.balance(id)); // the invariant, not "no errors"
}
Three details make the difference between a test that finds the race and one that passes by luck:
- One key, not many. Aim every worker at the same entity. A load test with random IDs is the reason the race was missed in the first place.
- A real database with realistic isolation. An in-memory fake or a mocked repository has no concurrency semantics. Use the same engine and isolation level as production, for example through a container in CI.
- A widened window. If the race still does not show, add a test-only hook between the read and the write (a short sleep, or a latch that holds the first worker until the second has read). A race that fails 1 run in 10,000 then fails nearly every run.
If a concurrent test passes most of the time and fails sometimes, do not mark it flaky and retry it. That intermittent failure is the race, caught cheaply. The cost of treating it as noise is an incident later.
Which fix fits which race condition?
Every correct fix does the same thing: it makes the read and the write indivisible, or it makes a conflicting write fail loudly instead of silently winning. Sleeps, retries, and “check again just before writing” only shrink the window. Shrinking the window lowers the frequency without removing the race.
Atomic conditional update
If the new value can be computed by the database from the old one, let the database do it in one statement. Put the condition in the WHERE clause and check how many rows changed:
const { rowCount } = await db.query(
`UPDATE accounts
SET balance = balance - $1
WHERE id = $2 AND balance >= $1`,
[amount, id]);
if (rowCount === 0) return res.status(409).send('insufficient funds'); // no stale read to act on
res.sendStatus(204);
This is safe at PostgreSQL’s default READ COMMITTED level. When two updates target the same row, the second waits for the first to commit. The PostgreSQL isolation documentation describes how the WHERE clause is then re-evaluated against the updated row. The second withdrawal sees the new balance, not the one it would have read earlier.
Optimistic locking with a version column
When the change needs application logic, such as a state machine transition or a document edit, read the row with its version and make the write conditional on the version being unchanged:
UPDATE orders
SET status = 'shipped', version = version + 1
WHERE id = $1 AND version = $2; -- $2 = the version you read
-- 0 rows updated: someone changed the order since you read it. Reload, re-decide, retry (bounded).
JPA and Hibernate do this for you with a @Version field. A conflicting write throws OptimisticLockException, which Spring translates to ObjectOptimisticLockingFailureException. Optimistic locking works well when conflicts are rare. On a hot row, most attempts conflict and the retries become their own load problem.
Pessimistic locking
SELECT ... FOR UPDATE locks the rows it reads until the transaction ends, so a second transaction waits instead of reading a value that is about to change. Keep the transaction short, and never make a remote call while holding the lock. The waiting requests each hold a database connection, which is how a lock turns into pool exhaustion. Lock rows in a consistent order to avoid deadlocks.
Stricter isolation
Under REPEATABLE READ or SERIALIZABLE, PostgreSQL aborts one of two conflicting transactions with could not serialize access due to concurrent update (or a serialization failure under SERIALIZABLE). That turns a silent lost update into an error, but only if the application retries the whole transaction. Raising the isolation level without adding that retry trades a data bug for an error rate.
A sleep, a retry, or a second check makes a race rarer. Only an atomic operation makes it impossible.
Races in memory: threads and async code
Not every race is in the database. Shared in-process state races too, and the tools differ by runtime.
Threads (Java, Go, and others)
count++ on a shared field is a read, an add, and a write, so concurrent threads lose increments. A plain HashMap modified from several threads can lose entries or corrupt its structure. Use the atomic operations the runtime provides: AtomicLong.incrementAndGet(), or ConcurrentHashMap.merge(key, 1L, Long::sum) instead of get followed by put. Go’s race detector (go test -race) and ThreadSanitizer find unsynchronized memory access at run time, and jcstress tests JVM concurrency behaviour. None of them can see a lost update between two requests that each talk to a database.
Async code (Node.js and friends)
“Node.js is single-threaded, so it has no races” is half true. Code between two awaits runs without interruption, but every await is a point where other requests run. A check before an await and an action after it is a check-then-act race, in one process with no threads:
let token = null;
async function getToken() {
if (!token || token.expiresAt < Date.now()) {
token = await fetchNewToken(); // 20 concurrent callers all see "expired" and all fetch
}
return token;
}
// Fixed: share the in-flight promise so only one refresh runs per process
let refreshing = null;
async function getToken() {
if (token && token.expiresAt >= Date.now()) return token;
refreshing ??= fetchNewToken().then(t => (token = t)).finally(() => { refreshing = null; });
return refreshing; // callers await the same refresh
}
The in-flight promise fixes the race inside one process. With several processes or replicas, each still refreshes once. That is fine for a token, but not for anything that must happen exactly once, which needs the database or another shared atomic primitive. The same pattern at cache scale is covered in Cache Stampede: Why a Healthy Cache Can Overload Your Database.
A mutex, synchronized block, or per-key promise map protects state within one process. Once the service runs on more than one replica, it does nothing for shared data in a database or cache. Distributed locks exist but are subtle: a lock holder can pause past its lease and keep writing. Prefer constraints and conditional writes in the data store itself.
What should you look for in review to catch races before production?
Races are cheapest to catch when the racy shape is introduced, because the shape is visible in code even when the timing is not. In review, look for:
- A read, then a decision or computation in code, then a write of the same entity
- An existence check followed by an insert
- An
awaitor remote call between a check and the action it guards - Shared mutable state in module scope, singletons, or static fields
- Is the endpoint retried by clients or the gateway?
- Is the job scheduled in app code on every replica?
- Does the consumer tolerate redelivery during processing?
- Which keys will be hot in production?
- Atomic statement, version check, lock, or constraint named
- Concurrent test on one key that asserts the invariant
- Conflict path handled: 409, retry with a bound, or reload
These checks belong to the broader practice in How to Review a Pull Request for Production Reliability Risks. The difference with races is that the reviewer must reason about two executions at once, which is exactly what a single-diff view makes hard.
How Tomosu helps
A race is visible in the code long before it is visible in the data. Tomosu analyzes the repository and each pull request for the shapes that turn into production races, and for the context a diff does not show:
- Read-modify-write and check-then-act paths: where an entity is read, decided on in application code, and written back without an atomic statement, version check, lock, or constraint.
- Windows that grow: remote calls, slow queries, and
awaits placed between the read and the write. - Hidden concurrency: endpoints behind retries, handlers for at-least-once messages, and jobs that run on every replica.
- Blast radius: which endpoints and consumers write the same entity, so a race on a payment balance is weighed above one on a display counter.
These findings feed the Production Reliability Index. Racy code gets discussed in review, with the concurrent test that would prove the fix, rather than reconstructed from inconsistent data weeks later.
Scan your repository with Tomosu →
Key takeaways
- A race needs two operations on the same data inside one read-to-write window. Overlaps scale with requests per key × window length.
- Production has hot keys, longer windows during slow moments, many replicas, and correlated arrivals. Tests usually have none of these.
- One user can create concurrency: retries after timeouts, double submits, redeliveries, and jobs running on every replica.
- Races rarely throw. Look for data inconsistent with the request history, and reconstruct access by entity ID.
- Reproduce with one key, a barrier start, many workers, many runs, a real database, and a widened window.
- Fix with an atomic conditional update, a version check, a row lock, or a constraint. Sleeps and retries only make races rarer.
- In-process locks and Node.js promise sharing do not protect data shared across replicas.
Frequently asked questions
Why do race conditions only appear under production load?
A race needs two operations on the same state inside the same short window. Tests send few requests, spread across unique keys, with fast dependencies, so overlaps almost never happen. Production has more concurrent requests, hot keys that receive a large share of traffic, retries, multiple replicas, and slow moments that widen the window.
How do I reproduce a race condition that only happens in production?
Aim many concurrent workers at the same entity, start them together with a barrier, and repeat the run many times. If it still does not fail, inject a delay between the read and the write in a test build to widen the window. Assert the invariant that broke in production, not just the absence of errors.
What is a lost update?
A lost update happens when two operations read the same value, each computes a new value from it, and both write. The second write overwrites the first, so one change disappears without any error. It is the most common race condition in web services that read a row, change it in application code, and write it back.
Can a race condition happen in single-threaded Node.js?
Yes. JavaScript code between two awaits runs without interruption, but every await lets other requests run. If a handler reads state, awaits something, and then writes based on what it read, another request can change the state in between. Multiple Node.js processes or replicas add true parallelism on top of that.
Should I use optimistic or pessimistic locking to fix a race?
Prefer a single atomic conditional update when the change fits in one statement. Use optimistic locking with a version column when conflicts are rare, and retry on conflict. Use pessimistic locking, such as SELECT FOR UPDATE, when the same row is contended often and retries would waste work, and keep that transaction short.
Do retries cause race conditions?
They often expose them. When a client times out and retries while the original request is still running on the server, one user action becomes two concurrent executions of the same operation. If that operation is a read-modify-write or a check-then-act, the retry can create a lost update or a duplicate.
Can the Go race detector or ThreadSanitizer find database race conditions?
No. They detect unsynchronized access to shared memory inside one process while the code runs. A lost update between two requests that each read and write a database row is not a memory data race, so it needs a concurrent integration test against a real database instead.
A race condition is two correct requests and one wrong result. Tomosu finds the read-to-write windows in your code before production traffic finds them for you. Assess your repository →