Streaming looks like the easy part. Call streamText, return a UI message
stream, tokens appear. Then production arrives:
- Your provider bill is higher than your traffic explains.
- Some answers stop mid-sentence and no error is ever logged.
- A user hits Stop and the UI clears, but the request keeps running.
- Under load, the app holds far more open connections than it has users.
All four are the same underlying fact: a streaming response is a long-lived connection with two ends that can each disappear independently, and the normal request/response error handling does not apply to it.
The code below is AI SDK 7. The toUIMessageStreamResponse() method on a
streamText result still works there, but it is deprecated in favour of the
stateless helpers shown here.
1. Errors after the first token do not reject
The moment the first chunk is written, the HTTP status is committed. There is no way to turn a 200 into a 500. So an error thrown mid-generation cannot propagate the way you expect: the response you returned already went out.
// The bug: this catch never fires for mid-stream failures.
try {
const result = streamText({ model, messages });
return createUIMessageStreamResponse({
stream: toUIMessageStream({ stream: result.stream }),
});
} catch (error) {
return Response.json({ error: "failed" }, { status: 500 });
}
streamText is not async: it returns immediately with a stream. The catch
covers only the synchronous setup: an unknown model id, a missing key. A
provider timeout at token 400 goes somewhere else entirely.
import { createUIMessageStreamResponse, streamText, toUIMessageStream } from "ai";
const result = streamText({
model: languageModel(),
instructions: systemPrompt(),
messages,
abortSignal: request.signal,
onError({ error }) {
// Mid-stream failures arrive here and nowhere else. Without this handler
// they are swallowed: the response just stops, with no log.
if (isAbortError(error)) return;
console.error("[chat] stream error", error);
},
});
return createUIMessageStreamResponse({
stream: toUIMessageStream({
stream: result.stream,
// What the *user* sees. Keep it generic: provider errors quote the prompt.
onError: () => "The assistant could not finish that response. Please try again.",
}),
});
Two different handlers, deliberately. onError on streamText is your log;
onError on toUIMessageStream is the user-visible string. Getting the
second one wrong is how a provider error message containing the system prompt
ends up in someone's browser.
2. The user leaves and you keep paying
When a browser tab closes or the user navigates away, the TCP connection drops. Nothing about that automatically stops the upstream generation: your server is still holding a request to Anthropic or OpenAI that will run to completion and bill you in full.
streamText({
model: languageModel(),
messages,
abortSignal: request.signal, // the code half of the fix
});
request.signal aborts when the client disconnects. Passing it through cancels
the upstream request, which stops both the tokens and the billing. It is one
line, it is easy to drop during a refactor, and its absence is invisible until
you read an invoice.
On Vercel that line is not enough on its own. Request cancellation is
opt-in per function. Without it the function keeps running after the client
leaves and request.signal never fires:
{
"$schema": "https://openapi.vercel.sh/vercel.json",
"functions": {
"src/app/api/chat/route.ts": { "supportsCancellation": true }
}
}
Two details make this bite. The path is the source file, with the src/
prefix when the app uses a src directory. And Vercel fails the deploy when a
listed path matches no function, so delete a route and its line together.
Cancellation also means the function can be stopped mid-work, so anything that
must finish after a disconnect (a usage row, a saved transcript) goes in
Next.js after().
The same signal covers the Stop button: stop() from useChat aborts the
fetch, which fires request.signal, which cancels upstream.
Two related details:
- Aborts are not errors. An
AbortErrorin your logs on every Stop click is noise that will train you to ignore the log. Filter it before logging: checkerror.namefor"AbortError"(and"ResponseAborted", which is what Next.js calls a dropped response). - Partial output may still be worth saving.
onEndruns only when a stream completes normally. An aborted stream callsonAbortinstead, with the steps that finished before the abort. If you persist conversations, save the partial turn there and mark it, rather than losing it.
3. What the user sees when it stops
The client half needs the same care:
const [transport] = useState(() => new DefaultChatTransport({ api: "/api/chat" }));
const { messages, sendMessage, status, stop, error, regenerate } = useChat({ transport });
const isBusy = status === "submitted" || status === "streaming";
status === "submitted": request sent, no tokens yet. This is where a "Thinking" indicator belongs. Without one, a slow first token reads as a broken button. Current models reason before the first visible token, so this window is longer than it used to be.status === "streaming": tokens arriving. Show Stop, not Send. A user who can only wait will reload, and a reload is a second billed request.error: the stream failed. Offerregenerate(). Do not silently drop the partial answer; leaving it visible with a retry beneath it is far less confusing than clearing the screen.- Disable submit while busy, or handle concurrent turns deliberately. The default double-send produces two overlapping streams into one message list.
Add aria-live="polite" to the message container. Without it a screen-reader
user gets silence while text streams in.
4. Backpressure and the platform
Backpressure is the case people invent complicated solutions for and then never hit, so be precise about what is real:
Real: each streamed response holds a connection and a function instance for its whole duration. Fifty concurrent thirty-second answers is fifty concurrent invocations. On a serverless platform that is a concurrency bill; on a single container it is a connection limit.
Real: maxDuration. The platform kills the function at the limit, and the
failure mode is a truncated stream with no error. With Vercel's Fluid compute
(the default for new projects) the default and the Hobby maximum are 300
seconds, and Pro goes to 800. Set it explicitly per route, sized to the work:
a chat turn that runs a minute is already a runaway.
export const runtime = "nodejs";
export const maxDuration = 60;
Real: buffering proxies. A CDN, an nginx in front of your app, or a
corporate proxy can buffer the whole response and deliver it at once: the user
waits the full duration and then sees everything. If streaming works locally
and not in production, this is the first thing to check. The SDK sends
x-vercel-ai-ui-message-stream: v1 and x-accel-buffering: no; make sure
nothing in front of you sets a Content-Length or buffers
text/event-stream.
Mostly not real: slow consumers. The Web Streams the SDK returns propagate backpressure correctly, and a browser reading tokens is never the bottleneck compared to token generation. You do not need to hand-tune this.
The checklist
For any route that streams:
- [ ]
abortSignal: request.signalpassed to the model call. - [ ] On Vercel, the route listed in
vercel.jsonwithsupportsCancellation: true. - [ ]
onErroronstreamText, logging server-side, skipping aborts. - [ ]
onErrorontoUIMessageStream, returning a generic user-facing string. - [ ]
maxDurationandmaxOutputTokensset to something the work fits inside. Reasoning tokens count againstmaxOutputTokens. - [ ] Client shows a distinct "submitted" state, a working Stop, and a retry.
- [ ] A finish reason other than
"stop"handled:"length"means you hitmaxOutputTokensand the answer is cut off mid-thought.
Test it by killing the tab mid-answer and watching your provider dashboard: the request should stop, not complete.