Someone opens an issue in Sentry to debug a failed checkout and finds the
customer's email, their delivery address, the full Cookie header, and (in
the request body) the raw contents of a form that included a phone number.
Nobody added any of that. Sentry collected it, because collecting request context is what an error tracker is for, and the defaults are tuned for debugging rather than for data protection.
Why the dashboard setting is not enough
Sentry has server-side data scrubbing, and it works. But look at where it runs: your process serialises the event, sends it over the network to Sentry's ingest endpoint, and then the scrubbing rules are applied before storage.
So by the time server-side scrubbing acts:
- the data has left your infrastructure and your jurisdiction
- it exists in transit logs and possibly in ingest buffers
- if your rules do not match the field, it is stored
The only scrubbing you fully control is the kind that happens in your process,
before the request is made. That is dataCollection, beforeSend,
beforeBreadcrumb and beforeSendSpan. Use the dashboard settings as a second
line, never as the first.
The wrong way
Sentry.init({
dsn: process.env.NEXT_PUBLIC_SENTRY_DSN,
// no dataCollection: Sentry 11's defaults collect cookies, bodies and more
beforeSend(event) {
if (event.user) delete event.user.email; // the one field you thought of
return event;
},
});
Two failures. Left at its defaults, Sentry 11's dataCollection attaches
cookies, request and response bodies, query strings, database values and local
variables, and fills in the user from the request. (Sentry 10 and earlier had a
sendDefaultPii flag for this; 11 replaced it with dataCollection.) And the
scrubber is an allowlist of one: it handles user.email and misses
request.data, breadcrumb payloads, tags, and any email that happens to be
inside an exception message.
The deeper problem is that the approach does not survive time. Someone adds a field to a form next quarter, and nothing about this code notices.
The right way: deny by default
Sentry.init({
dsn: process.env.NEXT_PUBLIC_SENTRY_DSN,
dataCollection: {
userInfo: false,
cookies: false,
httpBodies: [],
urlQueryParams: false,
databaseQueryData: false,
stackFrameVariables: false,
},
beforeSend: scrubErrorEvent,
beforeBreadcrumb: scrubBreadcrumb,
// Sentry 11 streams spans one by one; beforeSendTransaction no longer runs.
beforeSendSpan: scrubSpan,
});
with a scrubber built on four rules:
1. Drop request bodies. Do not filter them.
if (event.request) {
const { data: _data, cookies: _cookies, headers, query_string: _qs, url, ...rest } = event.request;
event.request = { ...rest, url: redactUrl(url), headers: redactHeaders(headers) };
}
A body is arbitrary user content, and "we redact the fields we thought of" stops being true the moment someone adds a field. If a body were genuinely needed to debug, you would want a specific, named, minimal projection of it, not the whole thing minus a blocklist.
2. Reduce the user to an opaque id.
event.user = event.user?.id ? { id: String(event.user.id) } : {};
The user field answers one question ("how many people are affected") and an opaque id answers it completely. An email answers it no better and turns your error tracker into a copy of your user table, in a third-party system, with a wider access list than your database.
3. Redact by value shape, everywhere.
Key-based rules miss anything inside a string. throw new Error(`No user for ${email}`)
puts an email in the issue title, where it is indexed and searchable, and no
field-level rule will catch it. So run patterns over every string in the event:
emails, long digit runs, bearer tokens, provider key prefixes, JWTs.
4. Drop console breadcrumbs entirely.
Breadcrumbs are collected automatically, and console breadcrumbs carry whatever
was logged. A single console.log(user) left in from a debugging session
becomes a user record attached to every error for the next hour.
if (breadcrumb.category === "console") return null;
You lose some debugging convenience. You gain the guarantee that a stray log statement cannot leak a record.
Do not forget the URL
URLs are the quiet leak. Query strings carry tokens (?reset_token=...),
emails (?email=...) and search terms people typed. Path segments carry ids.
Strip the query string entirely and replace id-shaped path segments:
/orders/9f3c1a2b-…/items?email=ada@example.com → /orders/:id/items
This has a second benefit: it improves grouping, because a hundred distinct URLs collapse into one route.
Make the scrubber total
It runs inside error handling. If it throws, you lose the error you were trying to report and the error from the scrubber is thrown inside the SDK, where it may be swallowed. So:
- depth-limit the walk (a deeply nested object should not recurse forever)
- guard against cycles with a
WeakSet - truncate long strings before running regexes over them
- no non-null assertions, no unguarded property access
Bound your regexes too. An unanchored greedy pattern run over every string of every event on a busy app is a measurable CPU cost in the hot path of your error handler.
What to keep
Scrubbing that removes everything makes the tracker useless, and people respond by turning it off. Keep what is diagnostic and impersonal:
- HTTP method, route pattern, status code
- an opaque user id, and a plan or role tag
- browser, OS, release, environment
- counts, durations, enums, booleans
- your own fingerprint hint
That set answers "what broke, for how many people, since which deploy", which is every question triage actually asks.
Verify with a real event
Configuration review is not verification. Trigger an error on a route that receives a realistic request, open the event, and read every section: Request, User, Tags, Contexts, Breadcrumbs, and the exception message itself. Do it again whenever you add a form, a header or a new integration.
Then write the patterns you added into a unit test, so the next refactor cannot quietly remove them.