The story is always the same. Someone ships an AI feature behind a form with no account required, posts it somewhere, and wakes up to a provider bill that is four figures larger than expected. The traffic was not organic. Someone found an endpoint that runs a language model, pointed a script at it, and your API key paid for their weekend.
There is no per-request cap on a model API. There is no "you have spent enough today" that fires by default. If a route can be called, it can be called ten thousand times an hour, and every one of those is billed to you.
Three signals that say "now"
Do not build metering into your prototype. Do build it before any of these is true:
- The route is reachable without authentication. This includes routes behind a session that anyone can create by signing up with a throwaway email. Rule of thumb: if getting a thousand requests costs an attacker nothing, the route is public.
- Any single request can cost more than a fraction of a cent. Long
context, tool loops, or a large
maxOutputTokensall multiply. A tool-using turn on a strong model with a big document in context is meaningfully expensive; a hundred of them is a bill you will notice. - You are about to price the product. The moment a plan says "500 messages a month", you need a counter that is the same one enforcement reads. Retrofitting the counter after selling the plan means reconciling two numbers that never agreed.
If none apply (an internal tool, five colleagues, a key with a $20 monthly cap on it), skip this. The provider dashboard's spend limit is enough.
Do the provider-side limit first
Before you write any code: set a monthly spend limit in the Anthropic or OpenAI console, and use a separate key per environment. This is a hard ceiling that works even when your application logic is wrong, and it takes two minutes. It is the only protection that cannot be bypassed by a bug in your own counting.
It is a ceiling, not a control: hitting it takes the feature down for everyone. Treat it as the last line, not the plan.
The wrong way
// Counts after the fact, in memory, per instance.
let requestsToday = 0;
export async function POST(request: Request) {
requestsToday += 1;
if (requestsToday > 1000) return new Response("Too many", { status: 429 });
// ...
}
Three problems. It is per-instance, so ten serverless instances give an attacker ten times the budget. It resets on every deploy and every cold start. And it counts requests, which is not the thing that costs money: one request with a 200k-token context costs more than a hundred short ones.
The right way, in three layers
Layer 1: a per-identity rate limit, before the model call
The cheapest effective control. It runs before you spend anything.
import { after } from "next/server";
export async function POST(request: Request) {
const identity = await identify(request); // user id, or hashed IP for anonymous
const gate = await rateLimit.check(identity, { limit: 20, windowSeconds: 3600 });
if (!gate.allowed) {
return Response.json(
{ error: "Rate limit reached. Try again in a few minutes." },
{ status: 429, headers: { "retry-after": String(gate.retryAfterSeconds) } },
);
}
// ...
}
Use a shared store: Redis, or your Postgres with a small rate_limit table
and an atomic upsert. Anything per-instance is decoration. Set the limit an
order of magnitude above what a real user does, not at it: the goal is to stop
scripts, not to annoy people.
For anonymous traffic, hash the IP with a server-side secret rather than storing it raw. That is a per-identity counter and not a personal data store.
Layer 2: record what each call actually cost
Every AI SDK result carries usage. Record it after the response is sent, so accounting never adds latency to the user's request.
const result = streamText({
model: languageModel(),
messages,
abortSignal: request.signal,
// AI SDK 7: `onEnd` (it was `onFinish`), and `usage` covers every step.
onEnd({ usage, finishReason, finalStep }) {
after(
recordUsage({
userId: identity,
feature: "chat",
model: finalStep.response.modelId,
inputTokens: usage.inputTokens ?? 0,
outputTokens: usage.outputTokens ?? 0,
finishReason,
}),
);
},
// An aborted stream skips `onEnd`. You were still billed for what ran.
onAbort({ steps }) {
after(
recordUsage({
userId: identity,
feature: "chat",
model: steps.at(-1)?.response.modelId ?? "unknown",
inputTokens: steps.reduce((sum, step) => sum + (step.usage.inputTokens ?? 0), 0),
outputTokens: steps.reduce((sum, step) => sum + (step.usage.outputTokens ?? 0), 0),
finishReason: "aborted",
}),
);
},
});
One row per call, with the model id the provider reports on it. Cost is derived at read time from a price table you keep in code. Do not store a computed cost: prices change and you will want to re-price history.
Record the aborted turn too. A step cut off mid-stream may not report its tokens at all, so a Stop click undercounts slightly; the provider dashboard is the source of truth for the invoice, your table is the source of truth for who spent it.
Layer 3: a quota check that reads the same numbers
Once layer 2 has data, the quota is a query:
const used = await tokensThisPeriod(userId);
const allowed = planLimits[plan].tokensPerMonth;
if (used >= allowed) {
return Response.json({ error: "Monthly limit reached.", upgradeUrl: "/pricing" }, { status: 402 });
}
Two details that decide whether people trust the number. Show it before they hit it: a counter in the UI at 80% turns a hard stop into an expected one. And decide deliberately whether a failed generation consumes quota; the honest answer is that you were billed, so it does, but say so in the plan description.
What to measure
Per feature, not just in total. "AI spend is up 40%" is not actionable; "the document-summary route is up 40% and its average input tokens tripled" is: somebody raised a truncation limit.
Keep these four:
- tokens in and out, per feature per day
- calls per user, so you can find the outlier
finishReasondistribution: a rise in"length"means answers are being cut off and users are re-asking, paying twice- p95 tokens per call, which catches context growing quietly
The cheapest wins, before you optimise anything
- Cap
maxOutputTokensper feature. Left unset, the SDK asks for the model's maximum (128,000 tokens on current Claude models). Leave room for reasoning, which counts against the cap, but a classification still does not need 4,000. - Use a smaller model where the task is easy. Routing, classification and
short extraction run fine on the cheap tier in
MODEL_CATALOGUE(Claude Haiku 4.5, GPT-6 Luna) at a fraction of the default model's cost. - Truncate context. Most "expensive" features are expensive because they send an entire document when the relevant paragraph would do.
- Prompt caching for a long, stable system prompt reused across requests: a provider feature that cuts repeat-context cost sharply.
- Batch anything not interactive. Both major providers offer a cheaper asynchronous tier for work nobody is waiting on.