Duplicate webhook events are not a provider bug. Webhooks are delivered at least once, so the same event will eventually reach your handler twice. The question is whether the second delivery ships a second parcel, sends a second transfer, and emails the customer again, or whether it becomes a harmless no-op.
To prevent duplicate webhook processing, make the first thing your handler does a write that can only succeed once: insert the provider’s event id into a table with a unique constraint, and only the delivery whose insert succeeds does the work. Then make every side effect idempotent on its own.
- Insert first, not check-then-act:
INSERT ... ON CONFLICT DO NOTHINGcloses the race aSELECTleaves open. - Acknowledge fast: store the event, return
2xx, process asynchronously. - Key every external effect: pass an idempotency key derived from the business entity, such as
transfer:ord_42. - Guard state transitions: events can arrive out of order, so never let a stale event overwrite newer state.
Most duplicate side effects come from a handler that looks correct in review. It verifies the signature, checks whether the event was seen, does the work, and records the event id. Every step is present, but they run in the wrong order and without a constraint, so two concurrent deliveries both pass the check. This guide covers why providers retry, where the idempotency boundary has to sit, and what an idempotent webhook handler looks like in code and SQL.
Why do duplicate webhook events happen?
A webhook provider cannot know whether your handler did the work. It only knows whether it received a successful HTTP response in time. When it does not, it has two choices: drop the event and risk you never learn a payment succeeded, or send it again and risk you process it twice. Providers that retry choose the second, which makes delivery at least once. Exactly-once delivery over an unreliable network is not on offer from anyone.
The situations that produce a second delivery of an event you already processed:
- Your handler was slow. It did the work but responded after the provider’s timeout. The provider records a failure and schedules a retry. GitHub, for example, expects a
2xxwithin 10 seconds. - Your handler returned a non-2xx after a partial success. The shipment was created, then the email call threw, and the framework returned a 500.
- The response was lost. A deploy, a load balancer reset, or a pod eviction dropped the connection after your code committed.
- Someone resent it. Dashboards and CLIs let an operator redeliver events by hand, usually during an incident when load is already unusual.
- The provider emitted it twice. Stripe’s documentation notes that in some cases two separate Event objects are generated for the same change, each with its own id.
Retry schedules differ by provider, and they are long enough that you cannot treat a duplicate as a rare edge case. The dedupe key differs too. These are the behaviors documented by the providers most teams integrate first:
| Provider | Dedupe key | Retry behavior | What to design for |
|---|---|---|---|
| Stripe | Event id (evt_...) | Automatic retries for up to three days with exponential backoff in live mode; manual resend from the Dashboard or CLI | Order is not guaranteed, and occasionally two Event objects describe one change, so also dedupe on data.object.id plus type |
| GitHub | X-GitHub-Delivery header | Expects a 2xx within 10 seconds; failed deliveries are not redelivered automatically, you redeliver them | The delivery GUID stays the same on redelivery, so it is a usable dedupe key |
| Standard Webhooks | webhook-id header | The spec recommends exponential backoff with jitter over multiple days | The id stays the same across retries of the same message |
Check your own provider’s documentation for its id field and retry window. The design below works regardless of the numbers.
Developer questions about duplicate webhook events on Stack Overflow tend to ask how to stop the provider from sending duplicates. You cannot, and you do not want to: the retries are what save you when your endpoint is down. The fix belongs in your handler.
Why checking the event id first is not enough
The first idempotent webhook handler most people write looks up the event id, returns early if it exists, does the work, and then records the id. Stripe’s own guidance, “log the event IDs you’ve processed, and then not process already-logged events,” reads naturally as exactly that. It handles the case where the duplicate arrives an hour later. It fails in the case that actually produces duplicates: a retry that arrives while the first delivery is still running, because the first delivery was slow. That is the timeout scenario from the diagram above.
SELECT cannot see a row another transaction has not written yet. A unique index can, and it makes the second writer wait for the first.Two details make insert-first work, and both are easy to get wrong:
- The constraint has to exist in the database. A
processed_eventstable without a primary key or unique index on the event id gives you a log, not a guard. The race is closed by the index, not by application code. - The insert has to happen before the work, in the same transaction as the work’s local writes. In PostgreSQL, when two transactions insert the same key into a unique index, the second one waits until the first commits or rolls back. If the first commits,
ON CONFLICT DO NOTHINGinserts nothing and the second learns it is a duplicate. If the first rolls back, the second proceeds and does the work. That is exactly the behavior you want when the first attempt crashes halfway.
Some handlers fix the race by inserting the event id in its own autocommitted statement and then doing the work. That closes the race and opens a worse hole: if the process dies after the insert and before the work, the retry is rejected as a duplicate and the event is never processed. Either commit the event id together with the state change, or give the row a status (pending, processing, done) and process it from there.
What an idempotent webhook handler looks like
There are two workable shapes. Choose based on what the event causes.
Shape 1: the event only changes your own database
If processing is a handful of writes to your own tables, do it synchronously in one transaction that starts with the dedupe insert. The event record and the state change commit together or not at all.
CREATE TABLE processed_events (
provider text NOT NULL,
event_id text NOT NULL,
processed_at timestamptz NOT NULL DEFAULT now(),
PRIMARY KEY (provider, event_id) -- the guard lives here
);
BEGIN;
INSERT INTO processed_events (provider, event_id)
VALUES ('stripe', $1)
ON CONFLICT DO NOTHING
RETURNING event_id;
-- no row returned: duplicate delivery. COMMIT and respond 200.
UPDATE orders SET status = 'paid', paid_at = now()
WHERE id = $2 AND status = 'pending_payment'; -- guarded transition
COMMIT;
This keeps a transaction open only for a few short statements. Do not stretch it around an HTTP call to another service: that holds a database connection for the length of the call, which is how a slow dependency turns into connection pool exhaustion.
Shape 2: the event causes external side effects
Charging, refunding, transferring money, creating a shipment, and sending an email cannot be rolled back by your database transaction. For these, split the handler in two. The receiver verifies the signature, stores the event in an inbox table under a unique constraint, and acknowledges. A worker processes the inbox. This is the pattern both Stripe and GitHub recommend when they tell you to respond quickly and process asynchronously.
CREATE TABLE webhook_inbox (
id bigserial PRIMARY KEY,
provider text NOT NULL,
event_id text NOT NULL,
event_type text NOT NULL,
payload jsonb NOT NULL,
-- status: pending | processing | done | failed
status text NOT NULL DEFAULT 'pending',
attempts int NOT NULL DEFAULT 0,
locked_until timestamptz,
received_at timestamptz NOT NULL DEFAULT now(),
UNIQUE (provider, event_id)
);
CREATE TABLE side_effects (
effect_key text PRIMARY KEY, -- e.g. 'receipt:ord_42'
created_at timestamptz NOT NULL DEFAULT now()
);
const raw = express.raw({ type: 'application/json' });
app.post('/webhooks/stripe', raw, async (req, res) => {
let event: Stripe.Event;
try {
// Raw body, not parsed JSON: the signature covers the exact bytes.
const sig = req.headers['stripe-signature'] as string;
event = stripe.webhooks.constructEvent(req.body, sig, endpointSecret);
} catch {
return res.sendStatus(400); // unsigned or tampered: record nothing
}
try {
await db.query(
`INSERT INTO webhook_inbox (provider, event_id, event_type, payload)
VALUES ('stripe', $1, $2, $3)
ON CONFLICT (provider, event_id) DO NOTHING`,
[event.id, event.type, event]);
} catch (err) {
return res.sendStatus(500); // not stored: let the provider retry
}
res.sendStatus(200); // stored now or earlier: acknowledge
});
Note what the receiver does not do: it does not look at the order, call any API, or branch on whether the insert found a conflict. A duplicate and a first delivery get the same fast 200. The only way to get a non-2xx is to fail to store the event, which is exactly when you want the provider to try again.
async function processNext(): Promise<boolean> {
const { rows: [job] } = await db.query(`
UPDATE webhook_inbox
SET status = 'processing', attempts = attempts + 1,
locked_until = now() + interval '5 minutes'
WHERE id = (SELECT id FROM webhook_inbox
WHERE status = 'pending'
OR (status = 'processing' AND locked_until < now())
ORDER BY id
FOR UPDATE SKIP LOCKED
LIMIT 1)
RETURNING *`);
if (!job) return false;
const session = job.payload.data.object as Stripe.Checkout.Session;
const order = await orders.findByCheckoutSession(session.id);
// Local state: a guarded transition is idempotent by construction.
await db.query(
`UPDATE orders SET status = 'paid'
WHERE id = $1 AND status = 'pending_payment'`, [order.id]);
// External effects: each key names the business fact, not the delivery.
await warehouse.createShipment(order,
{ idempotencyKey: `shipment:${order.id}` });
await stripe.transfers.create(
{ amount: order.sellerAmount, currency: 'usd',
destination: order.sellerAccountId, transfer_group: `order_${order.id}` },
{ idempotencyKey: `transfer:${order.id}` });
await sendOnce(`receipt:${order.id}`, () => mailer.sendReceipt(order));
await db.query(
`UPDATE webhook_inbox SET status = 'done' WHERE id = $1`, [job.id]);
return true;
}
The claim query uses FOR UPDATE SKIP LOCKED so several workers can drain the inbox without picking the same row, and locked_until lets another worker reclaim a row whose worker died mid-flight. Add a maximum attempt count that moves a row to failed and alerts; an event that fails forever should page someone rather than retry silently for days. The claim runs in its own short statement, so no transaction or connection is held while the worker waits on the warehouse or Stripe.
Because a reclaimed row runs the whole function again, every step after the claim must be safe to repeat. That is the job of the next section.
How do you stop a webhook retry from causing a duplicate payment?
The inbox stops duplicate deliveries. It cannot stop the worker from repeating an external call if the worker crashes between the call and marking the row done. A webhook retry duplicate payment almost always happens at this second layer: the dedupe was on the event, but the money movement had no key of its own.
Give every outgoing call that creates or moves something an idempotency key:
- Payment APIs that support it. Stripe accepts an
Idempotency-Keyheader onPOSTrequests (theidempotencyKeyrequest option in stripe-node). It saves the first result for a key and returns it for later requests with the same key, and it errors if the parameters differ. See Stripe’s idempotent requests reference. - Your own internal services. If the warehouse service does not accept an idempotency key, add one: a unique constraint on
(client_id, idempotency_key)on its side is the same pattern as the inbox. - Effects with no key at all, such as most email and SMS sends. Record the effect in a
side_effectstable under a unique key before sending. That makes the send at-most-once: a crash between the insert and the send loses the email. Recording after the send makes it at-least-once. Pick deliberately; for a receipt, a rare missing email usually beats a rare duplicate charge notice.
async function sendOnce(effectKey: string, send: () => Promise<void>) {
const { rowCount } = await db.query(
`INSERT INTO side_effects (effect_key) VALUES ($1)
ON CONFLICT DO NOTHING`, [effectKey]);
if (rowCount === 0) return; // already sent, or attempted
await send(); // at-most-once: a crash here loses this send
}
It is tempting to use transfer:${event.id}. Keying on order.id is stronger. If the provider emits two distinct events for the same change, or two different event types (a checkout completion and a payment success) both trigger fulfillment, an event-scoped key lets both through. The thing that must happen once is “pay the seller for order 42”, so that is what the key should name.
Also check the downstream key’s lifetime. Stripe may prune idempotency keys once they are at least 24 hours old, and a reused key after pruning is treated as a new request. A worker retrying a stuck row on day two is outside that window, so a local record of the effect (a transfer_id column on the order) is still worth having.
| Crash point | Without keys | With the inbox only | With inbox and effect keys |
|---|---|---|---|
| Before the inbox insert commits | Retry processes normally | Retry processes normally | Retry processes normally |
| Response lost after insert | Retry repeats everything | Retry is a no-op | Retry is a no-op |
Worker dies after transfer, before done | Second transfer | Second transfer on reclaim | Same key returns the original transfer |
| Provider emits two Event objects | Two transfers | Two transfers (different ids) | Keyed on order: one transfer |
What about webhooks that arrive out of order?
Retries reorder events. If event 1 fails its first delivery and is retried a minute later, event 2 arrives first. Stripe says plainly that it does not guarantee delivery order, and warns that its created timestamp has one second resolution, so it cannot be used to order events or to detect duplicates either.
Three techniques, usually combined:
- Guarded transitions. Write every status change as
UPDATE ... WHERE status IN (allowed previous states). Terminal states such aspaid,refunded, orcanceledare never overwritten by an event that implies an earlier state. - Re-fetch the current object. Treat the webhook as a notification that something changed, then retrieve the object from the provider’s API and act on its current state. Stripe’s docs suggest retrieving missing objects this way when an event arrives before the ones it depends on.
- Compare versions when the provider gives you one. If the payload carries a monotonic version or sequence number for the object, store it and ignore anything older. Do not invent ordering from delivery time.
How long to keep processed event ids
Longer than the longest window in which a duplicate can arrive. For Stripe that is the automatic retry period of up to three days plus manual resends, which its docs allow for 15 days from the Dashboard and 30 days from the CLI. A retention of 30 days or more with a scheduled cleanup is a reasonable default. A short TTL cache in front of the table is fine for load, but the database constraint should remain the source of truth.
Where signature verification fits
Verify the signature first, against the raw request body, and reject failures before recording anything. Signature checks stop forged and tampered requests, and the timestamp tolerance limits replay of old captures. They do nothing about duplicates: a legitimate retry is correctly signed. Stripe generates a new signature and timestamp for each delivery attempt, which is why you cannot dedupe on the signature header either.
Reviewing a PR that adds a webhook handler
Here is a PR of the kind that passes review in most teams. It adds a Stripe checkout.session.completed handler for a marketplace: ship the order, pay the seller, email a receipt. It has a signature check, it has a dedupe check, and it has tests that send an event once and assert the order ships.
app.post('/webhooks/stripe', raw, async (req, res) => {
let event: Stripe.Event;
try {
const sig = req.headers['stripe-signature'] as string;
event = stripe.webhooks.constructEvent(req.body, sig, endpointSecret);
} catch { return res.sendStatus(400); }
if (event.type !== 'checkout.session.completed') return res.sendStatus(200);
const seen = await db.query(
'SELECT 1 FROM processed_events WHERE event_id = $1', [event.id]);
if (seen.rowCount) return res.sendStatus(200); // check ...
const session = event.data.object as Stripe.Checkout.Session;
const order = await orders.findByCheckoutSession(session.id);
await warehouse.createShipment(order); // no key
await stripe.transfers.create({ // moves money, no key
amount: order.sellerAmount, currency: 'usd',
destination: order.sellerAccountId });
await mailer.sendReceipt(order); // slow: can exceed timeout
await db.query('UPDATE orders SET status = $2 WHERE id = $1',
[order.id, 'paid']); // unguarded
await db.query('INSERT INTO processed_events (event_id) VALUES ($1)',
[event.id]); // ... then act
res.sendStatus(200);
});
// migration.sql
// CREATE TABLE processed_events (event_id text, processed_at timestamptz);
// no primary key, no unique index
Each line is reasonable in isolation. The problem is where the idempotency boundary sits relative to the side effects:
The review findings, in order of consequence:
- Check-then-act with no constraint. The
SELECTcannot see a concurrent delivery, and the migration has no unique index, so even the lateINSERTcannot fail on a duplicate. - Side effects before the boundary. The transfer moves money on every delivery that passes the check. This is the webhook retry duplicate payment, waiting for the first slow SMTP response.
- Synchronous work that invites the retry. Three network calls inside the request make a provider timeout likely, and the timeout is what produces the concurrent duplicate.
- No idempotency keys on outgoing calls. Even after the boundary is fixed, a worker retry would repeat the transfer and the shipment.
- Unguarded state update. A stale or reordered event can overwrite
paid. - Tests that deliver once. None of this is visible until a test sends the same event twice, concurrently.
A useful test to require in the PR: fire the same signed payload at the endpoint from two concurrent requests, then assert one shipment, one transfer, and one email. It fails against the original handler and passes against the inbox version. For how this fits a wider review routine, see Pre Merge Reliability Analysis and How to Identify High Risk Pull Requests.
An idempotent handler is not one that checks for duplicates. It is one where nothing irreversible happens before a write that can only succeed once.
How Tomosu helps
The PR above is hard to catch in line-by-line review because every individual line is defensible. The risk is in the ordering and in what the called functions do: createShipment and transfers.create look like any other await. Tomosu analyzes the change against the code paths it touches and flags this class of problem before merge:
- The idempotency boundary. For a new or changed webhook or message handler, Tomosu maps where the event is first persisted and which calls run before it, and flags payments, shipments, and notifications that sit on the wrong side.
- Check-then-act dedupe. A
SELECTused as a guard with no matching unique constraint in the migrations is flagged as a race. - Outgoing calls without idempotency keys. Calls that create or move something inside a handler that can be retried are surfaced for review.
- Blast radius. A duplicate audit log line and a duplicate seller payout are weighted very differently. Findings are weighed by what the side effect touches, so the money path rises to the top. See How to Assess the Blast Radius of a Code Change.
These findings feed the Production Reliability Index, so a webhook PR with a missing idempotency boundary shows up as elevated risk at review time rather than as a double payout in the next retry storm.
Scan your repository with Tomosu →
Key takeaways
- Webhooks are delivered at least once. Duplicate webhook events are normal, and slow handlers make them concurrent.
- Check-then-act dedupe has a race. Insert the event id first under a unique constraint with
ON CONFLICT DO NOTHING. - Commit the event id with the state change, or give it a status. An id recorded on its own can mark an unprocessed event as done.
- Store, acknowledge with a
2xx, and process asynchronously. Only fail the request when the event was not stored. - Give every external side effect its own idempotency key, named after the business fact, and check how long the downstream keeps it.
- Events arrive out of order. Guard state transitions or re-fetch the current object; never trust the last delivery.
Frequently asked questions
Why do I receive duplicate webhook events?
Webhook providers deliver at least once. If your endpoint does not return a 2xx response in time, returns an error after partly processing the event, or the response is lost on the network, the provider retries with the same event id. Operators can also resend events manually, and some providers occasionally emit two separate events for one change. Duplicates are expected behavior, not a provider bug.
How do I make a webhook handler idempotent?
Make the first write one that can only succeed once: insert the provider’s event id into a table with a unique constraint using INSERT ... ON CONFLICT DO NOTHING, and only the delivery whose insert succeeds does the work. Commit that insert together with your state changes, or store the event with a status and process it from a queue. Then give every external side effect its own idempotency key.
Is checking whether the event id exists before processing enough?
No. A SELECT followed by processing and then an INSERT is a check-then-act race. Two concurrent deliveries, which is what a provider timeout produces, can both see no row and both run the side effects. The dedupe must be enforced by a unique constraint and the insert must happen before the work.
Should I return 200 before processing the webhook?
Return 2xx after the event is durably stored, not after it is fully processed. Verify the signature, insert the event into an inbox table, respond, and process it asynchronously. Return a non-2xx only when you failed to store the event, so the provider retries in exactly the case where you need it to.
How do I prevent a duplicate payment when a webhook is retried?
Deduplicate the delivery with a unique event id, and also send an idempotency key on the payment call itself, for example the Idempotency-Key header on Stripe POST requests. Derive the key from the business fact, such as the order id, rather than from the event id, and keep a local record of the created payment in case the retry happens after the provider has pruned the key.
How do I handle webhooks that arrive out of order?
Do not rely on delivery order or on event timestamps. Write status changes as guarded transitions that only apply from allowed previous states, never overwrite terminal states such as paid or refunded, and when in doubt retrieve the object’s current state from the provider’s API before acting.
Does webhook signature verification prevent duplicate processing?
No. Signature verification proves the request came from the provider and was not modified, and a timestamp tolerance limits replay of old requests. A legitimate retry is correctly signed, and providers such as Stripe generate a new signature for each delivery attempt, so you still need event id deduplication.
How long should I store processed webhook event ids?
Longer than the longest window in which a duplicate can arrive, including the provider’s automatic retry period and any manual resend window. For Stripe that means at least several days of automatic retries plus manual resends allowed for weeks, so 30 days or more with scheduled cleanup is a reasonable default.
A duplicate charge is rarely a missing check. It is a side effect that ran before the one write that could only succeed once. Tomosu maps where that write sits in every handler you ship. Assess your repository →