A customer pays once and gets two months of credit. Someone receives the welcome email three times. A subscription that was cancelled shows as active. All three have the same root cause, and it is not a bug in your handler logic.
Stripe guarantees at-least-once delivery. Not exactly-once. The same event arrives twice because:
- your endpoint took longer than Stripe's timeout, so it retried, and your first invocation also finished;
- your endpoint returned a 500 after doing half the work;
- two serverless instances were handed the same POST;
- someone clicked "Resend" in the dashboard;
- Stripe retried a delivery it could not confirm, for up to three days.
And events arrive out of order. A customer.subscription.updated generated
200 ms ago can land after the customer.subscription.deleted that superseded it.
The naive handler
// don't
export async function POST(request: Request) {
const event = JSON.parse(await request.text());
if (event.type === "checkout.session.completed") {
await grantCredits(event.data.object.client_reference_id, 100);
await sendWelcomeEmail(event.data.object.customer_email);
}
return new Response("ok");
}
Three defects: no signature verification (anyone can POST this and grant themselves credits), no idempotency (a retry grants 100 more), and it trusts the payload's contents rather than re-reading from Stripe.
Defence 1: verify the signature, on the raw bytes
const payload = await request.text(); // the exact bytes Stripe signed
const signature = request.headers.get("stripe-signature");
const event = await stripe.webhooks.constructEventAsync(
payload,
signature ?? "",
process.env.STRIPE_WEBHOOK_SECRET ?? "",
);
await request.json() cannot be used here: re-serialising reorders keys and
reformats numbers, and verification fails for every event. This is also why no
body-parsing middleware may sit in front of the route.
Use constructEventAsync, not constructEvent, wherever Web Crypto is the only
crypto available (edge runtimes, some serverless environments).
Defence 2: claim the event id before running the handler
An events table with the Stripe event id as its primary key:
create table billing_processed_events (
id text primary key, -- evt_...
type text not null,
claimed_at timestamptz not null default now(),
completed_at timestamptz
);
-- the claim itself, in one statement
insert into billing_processed_events (id, type) values ($1, $2)
on conflict (id) do nothing
returning id;
Wrapped in the three operations the handler actually calls. In a generated
repo they live in src/lib/billing/store.ts (written once per ORM, so nothing
else in the billing tree names a table), and processWebhook in
src/lib/billing/webhook.ts calls them. The key is prefixed with the provider,
so the same table serves whichever payments provider the repo uses:
const key = `stripe:${event.id}`;
if (!(await claimEvent(key, event.type))) return outcome("duplicate");
try {
await applyBillingEvents(await provider.translate(event));
} catch (error) {
// Release the claim so Stripe's retry is not swallowed as a duplicate.
await releaseEvent(key);
return new Response("Handler failed", { status: 500 });
}
await completeEvent(key);
Three details carry the weight:
Claim before, not after. Checking "have I seen this?" and later inserting leaves a window where two concurrent copies both pass the check. The insert is the check: the primary key does the mutual exclusion, in the database, atomically.
Release on failure. If the handler throws, the row is deleted so Stripe's retry can pick the work up. Leaving the claim in place turns a transient failure into a permanently dropped event.
completed_at is the audit trail. A row with completed_at null and an old
claimed_at is a handler that died mid-flight. Stripe got no 2xx and retries,
and a retry more than 15 minutes after the claim takes it over and redoes the
work. The rows no retry reached are your reconciliation worklist, and
bun run billing:prune-events lists them for you on every pass.
The ledger is not permanent storage. It answers "have I run this event",
and that question expires once Stripe's three-day retry window closes. Nothing
else deletes from the table, so schedule billing:prune-events. It removes
completed rows past a retention window and never touches an incomplete one.
Defence 3: make handlers converge, not accumulate
Even with a ledger, write handlers that would be safe if they ran twice.
// Re-read from Stripe rather than trusting a possibly stale payload, then hand
// back the state. The core upserts it by id. Out-of-order delivery stops mattering.
case "customer.subscription.updated":
return subscriptionEvents(event.data.object.id); // retrieve, then map
In a generated repo that is translate in src/lib/billing/provider.ts. It
never writes: it returns BillingEvents, and applyBillingEvents upserts them.
One-time purchases work the same way. A refund re-reads the Checkout Session
behind the charge and rewrites the whole purchase row, and the upsert only lets
the status move forward (paid to refunded, never back), so a late
checkout.session.completed cannot undo a refund.
This is the single most valuable pattern on this page. Because the handler
re-fetches current state and upserts, an event that arrives late writes the same
thing an event that arrived on time would have written. It also makes recovery
trivial: after an outage bun run billing:reconcile loops over
stripe.subscriptions.list() and converges on correct state without replaying
anything.
Re-reading has a second payoff: it pins the shape. A payload is rendered in the
API version of the endpoint that sent it, which can be older than the version
the SDK is pinned to. On a pre-Basil endpoint an invoice still carries
invoice.subscription and has no invoice.parent, so code that reads the new
field from the payload finds nothing and silently skips the work. The invoice
handlers here call stripe.invoices.retrieve(invoice.id) first for exactly that
reason, and the id is the one field every version agrees on.
The operations that genuinely accumulate (sending an email, granting credits, posting to Slack) are the ones the ledger protects. Keep them behind it and, where you can, give them their own idempotency key too.
There is one rule for those side effects that is easy to get backwards. They
belong inside the claimed window, and they must not be able to fail the
handler. A receipt email is the canonical example: by the time it is sent, the
payment row is already written, so letting the mail provider's outage throw
would release the claim, answer Stripe with a 500, and buy you a redelivery
that redoes the whole event and mails the customer twice once mail recovers.
Catch it, log it, carry on. A missing receipt is a support ticket; a
rolled-back payment record is an outage. In this repo receipts are
payment.receipt events that applyBillingEvents sends inside the claim, with
the send wrapped in a try/catch.
Defence 4: answer fast, work later
Stripe's delivery timeout is short. If your handler charges through a PDF render or a third-party API, you will time out, Stripe will retry, and you will do the slow thing twice.
Claim the event, write the state change, return 200, and queue anything slow. "Return 200 quickly" is not a performance nicety; it is what keeps the retry machinery from amplifying your load.
Status codes, precisely
- 2xx: Stripe considers it delivered and never retries it.
- Anything else: Stripe retries with exponential backoff for up to three days, then disables the endpoint and emails you.
So: return 2xx for events you do not handle (acknowledge and ignore), and return non-2xx only when a retry would genuinely help. A malformed event you can never process should be logged and acknowledged, not retried 40 times.
Signature verification failure is the exception: return 400. That request did not come from Stripe, or your secret is wrong, and either way a retry is correct behaviour.
Testing it
bun run stripe:listen
# complete a checkout, note the evt_ id, then:
stripe events resend evt_xxx
Expect a 200, {"outcome":"duplicate"}, and no second row anywhere. Resending
is the only test that proves the guard exists, and it is the test that fails the first
time someone adds an email to a handler without thinking about delivery
semantics.