Skip to content

When AI work has to move to a background job, and what breaks if you wait

Serverless functions have a hard timeout that no amount of streaming avoids. The four signals that mean the work no longer belongs in the request, and the smallest queue that fixes it.

AI bundle5 min readships at docs/solutions/ai-bundle/when-to-move-ai-work-to-a-background-job.md

Tags: ai · background-jobs · queues · timeouts · serverless

Everything in a generated AI bundle runs inside the HTTP request. That is the right default: it is simple, it streams, and there is nothing to operate. It also has a hard ceiling, and the way you find the ceiling is a bug report that says "it works for small files".

The wall

A serverless function is killed at maxDuration. On Vercel with Fluid compute (the default for new projects) that is 300 seconds by default, 300 at most on Hobby and 800 on Pro. The number moves; there is always a number, and when you hit it the function is terminated mid-work. The user sees a truncated stream or a network error. Nothing is retried. Nothing is saved. The tokens you generated are billed anyway.

Streaming does not help. Streaming changes when the first byte arrives, not how long the function runs. A twelve-minute batch job that streams progress is still a twelve-minute function.

Four signals it is time

  1. The work can exceed the timeout for realistic input. Not average input: the 95th percentile. If summarising a 3-page PDF takes 8 seconds, a 200-page one takes eight minutes and will be uploaded on day one.
  2. Nobody is watching. Nightly enrichment, re-embedding a corpus, scoring every new signup. If no human is waiting for the tokens, holding a request open is pure cost.
  3. It has to survive failure. A provider 529, a deploy mid-run, an instance recycled. Inside a request, all three lose the work silently. A job can be retried from where it stopped.
  4. It fans out. "Process these 400 rows" is 400 model calls. In a request they run in one function against one rate limit; as jobs they run at a concurrency you control, with per-item retries.

If none of these hold (a chat turn, a single classification, an extraction from one pasted email), keep it in the request. A queue you did not need is infrastructure you now maintain.

The wrong way

export async function POST(request: Request) {
  const { documentIds } = await request.json();

  const summaries = [];
  for (const id of documentIds) {
    const doc = await loadDocument(id);
    summaries.push(await summarise(doc));   // 6s each, 40 documents
  }

  return Response.json({ summaries });
}

This dies at document seven. The first six were generated and billed, and nothing recorded them. The user retries, and you pay for those six again. Add Promise.all and it dies faster, now with provider 429s in the log.

Two variants of the same mistake are worth naming:

  • Fire-and-forget without a runtime that waits. Calling an async function and returning immediately means the platform freezes the instance as soon as the response is sent. The work stops mid-flight. Next.js has after() for short post-response work; it is bounded by the same function lifetime and is not a job queue.
  • A cron route that does everything. A single /api/cron/process-all that loops over the backlog hits the same wall, just at 3am where nobody sees it.

The right way

Step 1: make the unit of work one item

The first change is not the queue, it is the shape. One job = one document, one row, one model call, with a status you can read:

model SummaryJob {
  id         String   @id @default(cuid())
  documentId String
  status     JobStatus @default(QUEUED)   // QUEUED | RUNNING | DONE | FAILED
  attempts   Int      @default(0)
  result     String?
  error      String?
  createdAt  DateTime @default(now())
  updatedAt  DateTime @updatedAt

  @@unique([documentId])
  @@index([status, createdAt])
}

The unique constraint on documentId is the idempotency key: enqueueing twice does not process twice. attempts is what stops a poisoned input retrying forever.

Step 2: accept fast, process later

export async function POST(request: Request) {
  const { documentIds } = await request.json();

  await db.summaryJob.createMany({
    data: documentIds.map((documentId) => ({ documentId })),
    skipDuplicates: true,
  });

  return Response.json({ queued: documentIds.length }, { status: 202 });
}

202 with a job id, not 200 with a result. The client polls a status route, or you push over websockets. This single change removes the timeout from the user-facing path entirely.

Step 3: pick the smallest runner that fits

In rough order of how much you take on:

  • A hosted job runner (Inngest, Trigger.dev, QStash). Fastest path on serverless: they call your endpoint, handle retries with backoff, give you a dashboard and step-level durability. You write a function, not a worker.
  • Your platform's cron hitting a route that claims and processes a small batch (say ten items) and returns. Runs every minute. Crude, free, and perfectly adequate for a backlog that is not latency-sensitive. Claim rows with an atomic update so two overlapping runs cannot take the same item.
  • A real worker process on a container host, pulling from a queue. Right when volume is steady and high; wrong as a first step, because it is a second deployable to operate.

Step 4: make each job safe to run twice

Every job runner retries. Assume at-least-once delivery:

  • Claim the row atomically (update ... where status = 'QUEUED' returning the row) so two runners cannot both take it.
  • Write results with the job id as the key, so a duplicate write is a no-op.
  • Cap attempts. After three, mark FAILED with the error and stop. A malformed PDF should not be retried a thousand times at model prices.
  • Record token usage per job. This is where cost visibility comes from once the work is no longer in a request you can watch.

The middle ground

Not everything needs a queue. Two patterns sit between:

  • after() for short trailing work. Logging usage, writing an audit row, invalidating a cache. Runs after the response, still inside the function's lifetime. Do not put a model call in it that could exceed the remaining time.
  • Chunking in the client. For a list of twenty documents, have the browser send twenty requests with a small concurrency limit. Each is short, progress is natural, and failures are per-item. It is not durable (a closed tab stops it), but it turns a timeout into a progress bar in an afternoon.

How to know it worked

  • Kill a deploy mid-batch. The queued items should still be processed afterwards.
  • Enqueue the same document twice; assert one result row and one model call.
  • Feed it something that always fails and confirm it stops after attempts reaches the cap, with the error recorded and visible.