Skip to content

Clerk user-sync webhooks: duplicates, retries and out-of-order events

Svix retries and can deliver twice, and updates can arrive before creates. Make the write idempotent on the user id and reject stale payloads by timestamp.

Clerk5 min readships at docs/solutions/clerk/webhook-idempotency-and-ordering.md

Tags: clerk · webhooks · svix · idempotency · user-sync

Your user sync works in development. In production you eventually find one of these:

  • Two rows for the same person, or a unique-constraint violation in the logs every few hours.
  • A user whose name reverted to what it was three edits ago.
  • A user who exists in Clerk and not in your database, with a 200 in the delivery log so nothing retried.

All three come from the same misunderstanding: treating webhook deliveries as an ordered, exactly-once stream. They are neither.

What the delivery system actually guarantees

Clerk delivers through Svix, which promises at least once. Concretely:

  • Retries on failure. Any non-2xx response, or a timeout, is retried with backoff over several hours. Your handler will see the same event again.
  • Duplicates without failure. A response that was sent but not recorded (a connection reset after your write committed) looks like a failure to Svix and is retried anyway.
  • No ordering. user.created and user.updated can be processed concurrently by two serverless instances. The one that started first is not the one that commits first.

So the handler must be safe to run twice with the same event, and safe to run with an event older than one it has already applied.

Idempotency: upsert on the stable id

The Clerk user id is stable for the lifetime of the account. It is the natural key for the mirror.

create unique index users_clerk_user_id_key on users (clerk_user_id);
await db
  .insert(users)
  .values({ clerkUserId, email, firstName, lastName, role, clerkUpdatedAt })
  .onConflictDoUpdate({
    target: users.clerkUserId,
    set: { email, firstName, lastName, role, clerkUpdatedAt },
  });

Note what this is not: a select followed by an insert or update. That pattern has a window between the two statements, and two concurrent deliveries both find no row and both insert. The upsert pushes the race into the database, where the unique index resolves it atomically.

The unique index is doing the real work. Without it, onConflictDoUpdate has no conflict to detect and quietly inserts a duplicate.

Ordering: reject stale payloads

Idempotency alone still lets an older event overwrite a newer one. That is the "name reverted" bug: a retry of user.created from four minutes ago lands after user.updated.

Every Clerk user payload carries updated_at in epoch milliseconds. Store it, and refuse to apply anything older:

.onConflictDoUpdate({
  target: users.clerkUserId,
  set: { /* ... */ },
  where: sql`${users.clerkUpdatedAt} < ${record.updatedAt}`,
})

Now a stale delivery is a no-op that still returns 200, which is exactly right: the event was handled, and handling it correctly meant doing nothing.

That guard cannot protect a row that is gone. A user.updated that failed before the account was deleted is retried for hours, lands after user.deleted, finds no row and inserts one: the deleted person is back. So a delete also leaves a tombstone, and every upsert checks it first. This battery keeps it in the delivery ledger under user.deleted:<clerk user id> (see deletionMarker in src/lib/auth/user-sync.ts), because Clerk never reuses a user id and a key no Svix id can take needs no table of its own.

Deduplicate on svix-id, in two layers

The svix-id header is the delivery's identity, and deduplicating on it is what stops the same event being applied twice. Where you keep that record decides whether the deduplication is real:

  • In-process memory is a cache, not a mechanism. Serverless gives you many instances and none of them share a Set. It kills the common case (a retry hitting the same warm instance seconds later) and answers "no" for everything else, which is the safe direction to be wrong in.
  • A table is the mechanism. One row per delivery you handled, keyed on svix_id as the primary key, so two instances racing the same retry lose on the unique constraint rather than on a race between two reads.

This battery ships both: clerk_webhook_deliveries is the table, and the Set in src/lib/auth/webhook-idempotency.ts sits in front of it to save a round trip on the warm path. Neither replaces the idempotent write above: a dedupe record that was never committed still has to leave the user row correct.

Mark a delivery handled only after the handler succeeded. Marking on arrival swallows the retry of a delivery that failed halfway through.

Events you deliberately ignore write no ledger row. Nothing happened, so a retry of one is another no-op, and a row per ignored event grows the table for nothing.

If you add a handler whose side effect genuinely cannot be repeated (charging a card, sending an email) write its record in the same transaction as the ledger row, so "did it happen" and "did we record that it happened" cannot disagree.

Status codes decide what happens next

SituationReturnWhy
Signature invalid400Not from Clerk. A retry will not help.
Missing Svix headers400Malformed.
Event type you ignore200Handled: by design.
Stale payload skipped200Handled.
Database unreachable500You want the retry.
Signing secret missing500Config error. Never treat as "skip verification".

The expensive mistake is returning 200 after a failed write. Svix marks it delivered, and the user is missing from your database permanently, with a clean delivery log to prove nothing went wrong.

Keep the handler fast

Svix times out; a slow handler is a failed handler, and a failed handler is retried, after your slow work already half-happened. Welcome emails, workspace provisioning and analytics calls go on a queue or a separate request. The webhook does one thing: write the mirror.

Backfilling

Users who existed before the webhook did never sent an event. Page through the Backend API once and run the same upsert. Because the write is idempotent, this is safe to run repeatedly and safe to run while live deliveries are arriving.

Checking your work

  • Replay the same delivery twice from the Clerk dashboard: one row, two 200s.
  • Replay a user.created after a user.updated: the newer data survives.
  • Point the app at a stopped database and replay: 500, and the dashboard shows a retry scheduled.
  • select clerk_user_id, count(*) from users group by 1 having count(*) > 1 returns nothing.