You wire up a tool, ship it, and a week later a customer has two refunds for
one order. The logs show one conversation, one user message, and two calls to
issueRefund with identical arguments about four seconds apart.
Nobody wrote a retry loop. The retry came from somewhere else, and the tool was not built to survive being called twice.
What actually retries
Three different mechanisms can produce a second call, and they have nothing to do with each other.
1. The AI SDK's maxRetries. This retries the HTTP request to the model
provider: a 429, a 500, a dropped connection. It defaults to 2, so a call can
be attempted three times. Crucially, it retries the request, and the request
is the whole conversation so far. If the previous step already ran your tool,
the tool is not re-run by this retry; the SDK is only re-asking the model.
2. The model itself. With stopWhen: isStepCount(n), the model sees the
tool result and decides what to do next. A tool that returns something the
model reads as failure ({ error: "timeout" }, an empty object, a thrown
error surfaced as text) is an invitation to try again. The model does not know
your tool has side effects. It knows it did not get an answer.
3. Your own infrastructure. A function that times out and is retried by the platform, a user who hits send twice, a client that resubmits.
The one that surprises people is (2). It looks like a retry loop nobody wrote, because the loop is a conversation.
The wrong way
export const issueRefund = tool({
description: "Refund an order.",
inputSchema: z.object({ orderRef: z.string(), amount: z.number() }),
execute: async ({ orderRef, amount }) => {
const charge = await payments.refund({ orderRef, amount });
await db.order.update({ where: { ref: orderRef }, data: { refundedAt: new Date() } });
await email.send({ template: "refund", orderRef });
return { ok: true, refundId: charge.id };
},
});
Three failure modes are baked in:
- If
email.sendthrows, the refund has already happened but the tool reports failure. The model apologises and tries again. The customer gets two refunds and one email. - If the whole thing succeeds but the response stream drops before the model reads the result, the next attempt starts from a conversation where the tool has no result, and calls it again.
- If the model is asked for "a refund for the last two orders" it may emit two
parallel tool calls with the same
orderRefbecause it misread the history.
The right way
Make the tool idempotent on a key the caller supplies. The AI SDK gives you
toolCallId in the execute options. It is stable for one tool call, including
across the SDK's own request retries.
export const issueRefund = tool({
description:
"Issue a refund for one order. Only call this after the customer has confirmed the amount.",
inputSchema: z.object({
orderRef: z.string().regex(/^ORD-[0-9]{6}$/),
amountCents: z.number().int().positive().max(50_000),
}),
// AI SDK 7: the user id arrives through `toolsContext`, checked against this
// schema, never as an argument the model fills in.
contextSchema: z.object({ userId: z.string() }),
execute: async ({ orderRef, amountCents }, { toolCallId, context: { userId } }) => {
const order = await db.order.findFirst({ where: { ref: orderRef, ownerId: userId } });
if (!order) return { ok: false, reason: "no such order on this account" };
if (order.refundedAt) return { ok: true, alreadyRefunded: true, refundId: order.refundId };
// The payment provider dedupes on this key: a second call with the same
// key returns the first result instead of moving money again.
const charge = await payments.refund({ orderRef, amountCents, idempotencyKey: toolCallId });
await db.order.update({ where: { id: order.id }, data: { refundedAt: new Date(), refundId: charge.id } });
// Side effects that are not part of the money movement happen after, and
// their failure does not fail the tool.
void email.send({ template: "refund", orderRef }).catch((error) => {
console.error("[refund] receipt email failed", { orderRef, message: String(error) });
});
return { ok: true, refundId: charge.id };
},
});
Four things changed, and each maps to one of the failure modes:
- An idempotency key. Every payment API supports one. So does any
INSERTwith a unique constraint on that key. Without it, "did this already happen?" is unanswerable. - A pre-check that returns success.
alreadyRefunded: truetells the model the goal is achieved. Returning an error here is what causes the second attempt. - Ordering by reversibility. The irreversible step happens once and first; the recoverable steps (email, analytics, cache invalidation) happen after and cannot fail the tool.
- Failures are data.
{ ok: false, reason }ends the loop cleanly. A thrown error, by contrast, propagates as a tool error the model reads as "try differently".
Deciding what a failure should look like
| Situation | Return | Why |
|---|---|---|
| Not found, no match, nothing to do | { ok: false, reason } | The model tells the user; no retry |
| Already done | { ok: true, already: true } | Ends the loop, no second side effect |
| Bad input that passed the schema | { ok: false, reason } | The model can ask the user for a correction |
| Upstream 5xx, likely transient | throw | The SDK's retry is the right handler |
| Upstream 4xx, will never succeed | { ok: false, reason } | Retrying wastes a step and a billed request |
The rule of thumb: throw only when a retry could plausibly help. Everything else is a result.
Bound the loop
streamText({
model: languageModel(),
tools: { issueRefund },
toolsContext: { issueRefund: { userId: session.user.id } },
stopWhen: isStepCount(6),
maxRetries: 2,
messages,
});
isStepCount (called stepCountIs before AI SDK 7) is the wall at the end of the conversation loop; maxRetries is
the transport budget for each individual request. Both are needed and they do
different jobs. Six steps is a reasonable default for search-and-answer; raise
it per surface for genuinely multi-step work, never globally.
Partial failure in a multi-tool step
A model can emit several tool calls in one step, and they run concurrently. If three succeed and one throws, the successful three have already happened, and there is no rollback: you are not in a transaction.
The practical mitigations, in order of how much they help:
- Make each tool individually idempotent (above). Then a re-run is harmless.
- Never expose two tools that must succeed together. If two things must be atomic, that is one tool wrapping one transaction.
- Return enough in each result for the model to describe what did happen. A user told "I refunded the order but could not cancel the subscription" is in a much better position than one told "something went wrong".
Checking your work
- Call the tool twice with the same
toolCallIdin a test and assert the side effect happened once. - Force the second, non-critical side effect to throw and assert the tool still reports success.
- Add an eval case that asks for the same action twice in one conversation and assert the second answer says it was already done.