Skip to content

Reconciling after a missed Stripe webhook

A customer paid and the app does not know. How to find the gap, replay or re-sync the affected subscriptions and one-time purchases, and build a reconciliation job so the next outage is boring.

Stripe6 min readships at docs/solutions/stripe/reconciling-a-missed-webhook.md

Tags: stripe · webhooks · reconciliation · incident-response · billing · recovery

"I paid but the app still says free." Every subscription product hears this eventually, and it almost always means one thing: a webhook did not reach you, or reached you and failed.

The good news is that Stripe is the source of truth and your database is a projection of it. If the projection is wrong, you rebuild it from the source. The work is knowing where to look and doing it without double-granting anything.

Why events go missing

  • Your endpoint was down or slow. Stripe retries with backoff for up to three days, then disables the endpoint and emails you. If nobody read that email, you are now missing everything since.
  • A deploy changed the signing secret and every delivery got a 400.
  • Auth middleware started matching /api/webhooks/**. Stripe cannot authenticate; the signature is the authentication. Any middleware in front of the route breaks it.
  • The handler threw on one bad event and the claim was never released.
  • The endpoint is subscribed to the wrong events. The classic is an app that handles checkout.session.completed but never subscribed to customer.subscription.updated, so it learns about the first payment and never about anything after. For one-time payments, the missing one is usually charge.refunded: the refund never reaches the app and the buyer keeps the lifetime plan.
  • Wrong mode. Live-mode events going to an endpoint configured in test mode, or the reverse.

Step 1: confirm the gap

Start from Stripe, not from your logs.

Dashboard → Developers → Webhooks → your endpoint. Look at the delivery log: failed attempts, pending retries, and whether the endpoint is disabled. This tells you when it started and whether it is still broken.

Dashboard → Developers → Events, filtered by type and date. This is the list of what should have reached you.

Then compare against your ledger:

-- events you claimed but never finished: handler crashed mid-flight
select id, type, claimed_at from billing_processed_events
where completed_at is null order by claimed_at desc;

-- the last event you successfully processed
select max(completed_at) from billing_processed_events;

The gap between max(completed_at) and now is your window.

Fix the cause before you replay. Replaying into a broken endpoint just regenerates the failures.

Step 2: replay, for a small number of events

For a handful of known events, resend them. Stripe re-delivers with the same evt_ id, so your idempotency ledger correctly treats an already-processed one as a duplicate. Replaying is safe.

Dashboard → Developers → Events → the event → Resend, or:

stripe events resend evt_xxx

For a broader window, the endpoint's page has a bulk resend for failed deliveries. Watch your logs as they land; a burst of replays is also a load test.

Step 3: re-sync, for anything larger

Replaying hundreds of events is slower and more fragile than simply asking Stripe what is true now. That is what scripts/billing/reconcile.ts does, and it ships with this repo:

bun run billing:reconcile

It walks stripe.subscriptions.list({ status: "all" }) (an async iterator that pages for you) and turns every subscription into the same BillingEvents the webhook produces, through eventsForSubscription in src/lib/billing/provider.ts, then applies them with applyBillingEvents, which upserts by id. Then it re-reads every completed payment-mode Checkout Session from the last 90 days (bun run billing:reconcile -- 365 looks further) the same way, which also catches a refund or a dispute a webhook missed. It sends no mail: receipts belong to the webhook that saw the payment. Running it twice changes nothing the second time. That property is what makes it safe to run during an incident, at 2am, without thinking hard.

It prints four numbers, and each means something different:

  • subscriptions reconciled: written into the projection.
  • one-time purchases reconciled: re-read and upserted. The purchase upsert only moves a status forward, so a sweep can mark a purchase refunded but never un-refund it.
  • skipped: subscriptions whose customer has no billing_customers row and no metadata.userId: created in the dashboard, imported, or belonging to a different environment sharing the Stripe account. Investigate them by hand rather than inventing a user; step 4 below is how.
  • cancelled: local rows that were entitling someone and that a complete listing did not mention. Stripe cancels subscriptions, it does not delete them, so a row Stripe has never heard of is one no future event will ever correct. If the listing fails part way, nothing is cancelled. The script cannot tell a subscription Stripe dropped from one it never reached, and guessing would revoke access from paying customers.

If you need to reconcile without that last step (during a migration, say, or while two environments share a Stripe account), comment out the markSubscriptionsCanceled call rather than narrowing the listing, which would make the diff wrong in the other direction.

Step 4: repair the customer mapping, if that is the gap

Sometimes the subscription is fine and the link is missing. Two handles let you recover it:

  • subscription.metadata.userId, set in subscription_data.metadata when checkout started.
  • checkout_session.metadata.userId, set to the user id on the session. Not client_reference_id: a Payment Link lets its buyer type that one.
  • For one-time purchases, payment_intent.metadata.userId, which checkout also sets, so a refund in the dashboard can still be traced to a user.

If both are absent for a customer, match on email as a last resort: carefully, manually, and never in an automated loop. Email is not a unique identifier for a person and merging two customers into one user is not reversible.

Step 5: make the next one boring

Alert on the ledger, not on Stripe. A scheduled check that max(completed_at) is within the last hour catches a broken endpoint faster than Stripe's three-day retry window and its "we disabled your endpoint" email.

select now() - max(completed_at) as since_last_event from billing_processed_events;

Run the reconciliation job on a schedule: bun run billing:reconcile, nightly, from a Vercel cron or a scheduled GitHub Action. It is idempotent, it is cheap, and it turns "we lost events" from an incident into a line in a log. This is the single highest-value thing on this page.

Alert on stuck claims. Rows with completed_at null and claimed_at older than a few minutes are handlers that died. They are also events Stripe will not retry if you returned a 2xx before crashing. bun run billing:prune-events counts them for you on every run, alongside trimming the completed rows. Schedule it next to the reconciler. It never deletes an incomplete row, because that would hand the next redelivery a clean slate to re-run its side effects on.

Check the endpoint's subscribed event list against HANDLED_EVENTS in src/lib/billing/provider.ts after any change to either. They drift silently.

What not to do

  • Do not grant access manually and move on. You have fixed one customer and left the cause in place. Find the gap first, fix the cause, then reconcile everyone.
  • Do not delete rows from billing_processed_events to force a reprocess. That is how a retry re-runs a side effect. If you need to reprocess, replay the event and let the ledger do its job, or re-sync, which does not need the ledger at all. (billing:prune-events is not an exception: it only removes rows whose handler completed, and only long after Stripe's retry window has closed.)
  • Do not trust the event payload over a fresh fetch. An event from two days ago describes the world two days ago. Every case in the Stripe adapter's translate re-reads the object by id for exactly this reason.
  • Do not build a parallel source of truth. Your table is a cache of Stripe. Every recovery procedure gets simpler if you keep believing that.