Skip to content

Scrubbing PII before it leaves your process, not after it reaches Sentry

Server-side scrubbing runs after the data has crossed the network. beforeSend runs in your process, and it is the only layer you fully control.

Sentry5 min readships at docs/solutions/sentry/scrubbing-pii-before-it-leaves-the-process.md

Tags: sentry · pii · privacy · gdpr · beforeSend · security

Someone opens an issue in Sentry to debug a failed checkout and finds the customer's email, their delivery address, the full Cookie header, and (in the request body) the raw contents of a form that included a phone number.

Nobody added any of that. Sentry collected it, because collecting request context is what an error tracker is for, and the defaults are tuned for debugging rather than for data protection.

Why the dashboard setting is not enough

Sentry has server-side data scrubbing, and it works. But look at where it runs: your process serialises the event, sends it over the network to Sentry's ingest endpoint, and then the scrubbing rules are applied before storage.

So by the time server-side scrubbing acts:

  • the data has left your infrastructure and your jurisdiction
  • it exists in transit logs and possibly in ingest buffers
  • if your rules do not match the field, it is stored

The only scrubbing you fully control is the kind that happens in your process, before the request is made. That is dataCollection, beforeSend, beforeBreadcrumb and beforeSendSpan. Use the dashboard settings as a second line, never as the first.

The wrong way

Sentry.init({
  dsn: process.env.NEXT_PUBLIC_SENTRY_DSN,
  // no dataCollection: Sentry 11's defaults collect cookies, bodies and more
  beforeSend(event) {
    if (event.user) delete event.user.email;   // the one field you thought of
    return event;
  },
});

Two failures. Left at its defaults, Sentry 11's dataCollection attaches cookies, request and response bodies, query strings, database values and local variables, and fills in the user from the request. (Sentry 10 and earlier had a sendDefaultPii flag for this; 11 replaced it with dataCollection.) And the scrubber is an allowlist of one: it handles user.email and misses request.data, breadcrumb payloads, tags, and any email that happens to be inside an exception message.

The deeper problem is that the approach does not survive time. Someone adds a field to a form next quarter, and nothing about this code notices.

The right way: deny by default

Sentry.init({
  dsn: process.env.NEXT_PUBLIC_SENTRY_DSN,
  dataCollection: {
    userInfo: false,
    cookies: false,
    httpBodies: [],
    urlQueryParams: false,
    databaseQueryData: false,
    stackFrameVariables: false,
  },
  beforeSend: scrubErrorEvent,
  beforeBreadcrumb: scrubBreadcrumb,
  // Sentry 11 streams spans one by one; beforeSendTransaction no longer runs.
  beforeSendSpan: scrubSpan,
});

with a scrubber built on four rules:

1. Drop request bodies. Do not filter them.

if (event.request) {
  const { data: _data, cookies: _cookies, headers, query_string: _qs, url, ...rest } = event.request;
  event.request = { ...rest, url: redactUrl(url), headers: redactHeaders(headers) };
}

A body is arbitrary user content, and "we redact the fields we thought of" stops being true the moment someone adds a field. If a body were genuinely needed to debug, you would want a specific, named, minimal projection of it, not the whole thing minus a blocklist.

2. Reduce the user to an opaque id.

event.user = event.user?.id ? { id: String(event.user.id) } : {};

The user field answers one question ("how many people are affected") and an opaque id answers it completely. An email answers it no better and turns your error tracker into a copy of your user table, in a third-party system, with a wider access list than your database.

3. Redact by value shape, everywhere.

Key-based rules miss anything inside a string. throw new Error(`No user for ${email}`) puts an email in the issue title, where it is indexed and searchable, and no field-level rule will catch it. So run patterns over every string in the event: emails, long digit runs, bearer tokens, provider key prefixes, JWTs.

4. Drop console breadcrumbs entirely.

Breadcrumbs are collected automatically, and console breadcrumbs carry whatever was logged. A single console.log(user) left in from a debugging session becomes a user record attached to every error for the next hour.

if (breadcrumb.category === "console") return null;

You lose some debugging convenience. You gain the guarantee that a stray log statement cannot leak a record.

Do not forget the URL

URLs are the quiet leak. Query strings carry tokens (?reset_token=...), emails (?email=...) and search terms people typed. Path segments carry ids.

Strip the query string entirely and replace id-shaped path segments:

/orders/9f3c1a2b-…/items?email=ada@example.com   →   /orders/:id/items

This has a second benefit: it improves grouping, because a hundred distinct URLs collapse into one route.

Make the scrubber total

It runs inside error handling. If it throws, you lose the error you were trying to report and the error from the scrubber is thrown inside the SDK, where it may be swallowed. So:

  • depth-limit the walk (a deeply nested object should not recurse forever)
  • guard against cycles with a WeakSet
  • truncate long strings before running regexes over them
  • no non-null assertions, no unguarded property access

Bound your regexes too. An unanchored greedy pattern run over every string of every event on a busy app is a measurable CPU cost in the hot path of your error handler.

What to keep

Scrubbing that removes everything makes the tracker useless, and people respond by turning it off. Keep what is diagnostic and impersonal:

  • HTTP method, route pattern, status code
  • an opaque user id, and a plan or role tag
  • browser, OS, release, environment
  • counts, durations, enums, booleans
  • your own fingerprint hint

That set answers "what broke, for how many people, since which deploy", which is every question triage actually asks.

Verify with a real event

Configuration review is not verification. Trigger an error on a route that receives a realistic request, open the event, and read every section: Request, User, Tags, Contexts, Breadcrumbs, and the exception message itself. Do it again whenever you add a form, a header or a new integration.

Then write the patterns you added into a unit test, so the next refactor cannot quietly remove them.