Skip to content

Orphaned objects in R2: lifecycle rules and the sweep they cannot do

Direct-to-bucket uploads leak objects nobody references. Lifecycle rules clean up a tmp/ prefix and abandoned multipart parts; owner-scoped orphans need a reconciliation job.

Cloudflare R24 min readships at docs/solutions/r2/cleaning-up-orphaned-uploads.md

Tags: r2 · storage · lifecycle · cleanup · orphans · multipart · cost

Six months after shipping uploads, someone looks at the R2 dashboard: 40 GB stored, and your database references maybe 12 GB of it. Nothing is broken. You are simply paying to store files nothing points at, and the gap grows every week.

This is the structural cost of direct-to-bucket uploads, and it is worth understanding before you try to fix it.

Where orphans come from

Direct upload is two independent steps: the browser PUTs the bytes to R2, then tells your server the key so it can be recorded on a row. Anything that interrupts the gap between them leaves an object with no owner.

  1. The abandoned upload. The user picks a file, the PUT succeeds, they close the tab before the server action runs. Bytes in the bucket, no row.
  2. The replaced file. A user uploads a new avatar. The row now points at the new key. The old object is still there: this one is entirely in your control and is the most common leak in practice.
  3. The deleted row. A record is deleted and nobody deleted its object. Cascade deletes in the database do not cascade into a bucket.
  4. The failed multipart. Parts from an incomplete multipart upload are stored and billed, do not appear in a ListObjectsV2 result, and are invisible in the dashboard's object count. This is the leak people never find by looking.
  5. Test and seed data. Every developer's local uploads, if they share a bucket.

What lifecycle rules can do

R2 lifecycle rules delete objects by age, with a prefix filter, and the filter matches from the start of the key. That single constraint decides everything about how you use them.

infra/r2/lifecycle.json, applied with bun run r2:lifecycle:

{
  "Rules": [
    { "ID": "expire-temp-uploads", "Status": "Enabled",
      "Filter": { "Prefix": "tmp/" }, "Expiration": { "Days": 1 } },
    { "ID": "abort-incomplete-multipart", "Status": "Enabled",
      "Filter": { "Prefix": "" },
      "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 } }
  ]
}

The second rule is the one to apply everywhere, today, on every bucket. Abandoned multipart parts are billed storage that nothing else on earth will ever clean up, and no amount of looking at your object listing will reveal them.

The first rule is a design you can opt into: write unconfirmed uploads under a top-level tmp/ prefix, and move (copy + delete) the object to its permanent key only when your server records it. Anything never confirmed evaporates in a day, for free, with no job to run.

What lifecycle rules cannot do

The keys in this repo are <ownerId>/<prefix>/<uuid>-<name>, with the owner first, because ownership is then a property of the key, checkable without a database round trip. The cost of that layout is that no prefix filter can match "uploads older than 30 days across all owners": the owner segment varies, and R2 filters from the start of the key.

So lifecycle handles tmp/ and multipart. Owner-scoped orphans need reconciliation, and reconciliation needs your database.

The sweep

List the bucket, ask the database which keys it knows about, delete the difference. Two guardrails matter more than the code: a grace period, so you never delete an object uploaded while the job was running, and a dry run, so the first version prints instead of deleting.

// scripts/r2/sweep-orphans.ts
import { DeleteObjectsCommand, ListObjectsV2Command } from "@aws-sdk/client-s3";

const GRACE_MS = 24 * 60 * 60 * 1000;

async function sweep({ dryRun = true } = {}) {
  const cutoff = Date.now() - GRACE_MS;
  const orphans: string[] = [];
  let token: string | undefined;

  do {
    const page = await s3().send(
      new ListObjectsV2Command({ Bucket: bucket(), ContinuationToken: token }),
    );

    const candidates = (page.Contents ?? []).filter(
      (object) => object.Key && (object.LastModified?.getTime() ?? 0) < cutoff,
    );

    // One query per page, not one per object. A thousand round trips is how a
    // cleanup job becomes the thing that takes the database down.
    const keys = candidates.map((object) => object.Key as string);
    const known = new Set(await referencedKeys(keys));

    orphans.push(...keys.filter((key) => !known.has(key)));
    token = page.NextContinuationToken;
  } while (token);

  console.log(`${orphans.length} orphaned object(s)`);
  if (dryRun) {
    for (const key of orphans.slice(0, 50)) console.log(`  would delete ${key}`);
    return;
  }

  // DeleteObjects takes up to 1000 keys per call: one class A operation
  // instead of a thousand.
  for (let i = 0; i < orphans.length; i += 1000) {
    await s3().send(
      new DeleteObjectsCommand({
        Bucket: bucket(),
        Delete: { Objects: orphans.slice(i, i + 1000).map((Key) => ({ Key })) },
      }),
    );
  }
}

referencedKeys() is the part only you can write, and it must cover every column in every table that stores a key. A sweeper that knows about user.avatarKey but not message.attachmentKey does not leak storage: it deletes your customers' attachments. Write it as one query per table, union the results, and add a test that fails when a new key column appears.

Run it dry for a week. Read the list. Only then pass dryRun: false, and even then, schedule it weekly rather than hourly: the risk profile of a bulk delete does not improve with frequency.

Cheaper than any of this: do not create orphans

Sweeping is the safety net. The fixes are upstream:

  • Delete the old object when you replace a reference. Three lines in the server action that sets the new key, and it removes the largest source.
  • Delete the object in the same code path that deletes the row. Object first, then the row: an orphaned object costs money, an orphaned row breaks a page.
  • Write unconfirmed uploads under tmp/ and let the lifecycle rule handle abandonment, so the sweeper only ever has real bugs to find.
  • Keep short signed-URL lifetimes. Fifteen minutes bounds how long an abandoned intent stays uploadable at all.

Knowing whether it is working

Track two numbers monthly: total bytes in the bucket, from ListObjectsV2 or the Cloudflare dashboard, and the sum of what your own tables reference. The gap is your orphan estimate. Watching it is how you notice a new leak in the week it appears rather than the year it becomes expensive.