Everything in a generated AI bundle runs inside the HTTP request. That is the right default: it is simple, it streams, and there is nothing to operate. It also has a hard ceiling, and the way you find the ceiling is a bug report that says "it works for small files".
The wall
A serverless function is killed at maxDuration. On Vercel with Fluid compute
(the default for new projects) that is 300 seconds by default, 300 at most on
Hobby and 800 on Pro. The number moves; there is always a number, and when you
hit it the function is terminated mid-work. The user sees a truncated stream or a network error. Nothing is
retried. Nothing is saved. The tokens you generated are billed anyway.
Streaming does not help. Streaming changes when the first byte arrives, not how long the function runs. A twelve-minute batch job that streams progress is still a twelve-minute function.
Four signals it is time
- The work can exceed the timeout for realistic input. Not average input: the 95th percentile. If summarising a 3-page PDF takes 8 seconds, a 200-page one takes eight minutes and will be uploaded on day one.
- Nobody is watching. Nightly enrichment, re-embedding a corpus, scoring every new signup. If no human is waiting for the tokens, holding a request open is pure cost.
- It has to survive failure. A provider 529, a deploy mid-run, an instance recycled. Inside a request, all three lose the work silently. A job can be retried from where it stopped.
- It fans out. "Process these 400 rows" is 400 model calls. In a request they run in one function against one rate limit; as jobs they run at a concurrency you control, with per-item retries.
If none of these hold (a chat turn, a single classification, an extraction from one pasted email), keep it in the request. A queue you did not need is infrastructure you now maintain.
The wrong way
export async function POST(request: Request) {
const { documentIds } = await request.json();
const summaries = [];
for (const id of documentIds) {
const doc = await loadDocument(id);
summaries.push(await summarise(doc)); // 6s each, 40 documents
}
return Response.json({ summaries });
}
This dies at document seven. The first six were generated and billed, and
nothing recorded them. The user retries, and you pay for those six again. Add
Promise.all and it dies faster, now with provider 429s in the log.
Two variants of the same mistake are worth naming:
- Fire-and-forget without a runtime that waits. Calling an async function
and returning immediately means the platform freezes the instance as soon as
the response is sent. The work stops mid-flight. Next.js has
after()for short post-response work; it is bounded by the same function lifetime and is not a job queue. - A cron route that does everything. A single
/api/cron/process-allthat loops over the backlog hits the same wall, just at 3am where nobody sees it.
The right way
Step 1: make the unit of work one item
The first change is not the queue, it is the shape. One job = one document, one row, one model call, with a status you can read:
model SummaryJob {
id String @id @default(cuid())
documentId String
status JobStatus @default(QUEUED) // QUEUED | RUNNING | DONE | FAILED
attempts Int @default(0)
result String?
error String?
createdAt DateTime @default(now())
updatedAt DateTime @updatedAt
@@unique([documentId])
@@index([status, createdAt])
}
The unique constraint on documentId is the idempotency key: enqueueing twice
does not process twice. attempts is what stops a poisoned input retrying
forever.
Step 2: accept fast, process later
export async function POST(request: Request) {
const { documentIds } = await request.json();
await db.summaryJob.createMany({
data: documentIds.map((documentId) => ({ documentId })),
skipDuplicates: true,
});
return Response.json({ queued: documentIds.length }, { status: 202 });
}
202 with a job id, not 200 with a result. The client polls a status route, or you push over websockets. This single change removes the timeout from the user-facing path entirely.
Step 3: pick the smallest runner that fits
In rough order of how much you take on:
- A hosted job runner (Inngest, Trigger.dev, QStash). Fastest path on serverless: they call your endpoint, handle retries with backoff, give you a dashboard and step-level durability. You write a function, not a worker.
- Your platform's cron hitting a route that claims and processes a small batch (say ten items) and returns. Runs every minute. Crude, free, and perfectly adequate for a backlog that is not latency-sensitive. Claim rows with an atomic update so two overlapping runs cannot take the same item.
- A real worker process on a container host, pulling from a queue. Right when volume is steady and high; wrong as a first step, because it is a second deployable to operate.
Step 4: make each job safe to run twice
Every job runner retries. Assume at-least-once delivery:
- Claim the row atomically (
update ... where status = 'QUEUED'returning the row) so two runners cannot both take it. - Write results with the job id as the key, so a duplicate write is a no-op.
- Cap
attempts. After three, markFAILEDwith the error and stop. A malformed PDF should not be retried a thousand times at model prices. - Record token usage per job. This is where cost visibility comes from once the work is no longer in a request you can watch.
The middle ground
Not everything needs a queue. Two patterns sit between:
after()for short trailing work. Logging usage, writing an audit row, invalidating a cache. Runs after the response, still inside the function's lifetime. Do not put a model call in it that could exceed the remaining time.- Chunking in the client. For a list of twenty documents, have the browser send twenty requests with a small concurrency limit. Each is short, progress is natural, and failures are per-item. It is not durable (a closed tab stops it), but it turns a timeout into a progress bar in an afternoon.
How to know it worked
- Kill a deploy mid-batch. The queued items should still be processed afterwards.
- Enqueue the same document twice; assert one result row and one model call.
- Feed it something that always fails and confirm it stops after
attemptsreaches the cap, with the error recorded and visible.