Skip to content

Tool retries and partial failures: what the AI SDK retries and what it does not

The SDK retries the request to the model, never your tool's side effects. A tool that half-succeeded and then got called again is where duplicate charges come from.

AI bundle5 min readships at docs/solutions/ai-bundle/tool-retries-and-partial-failures.md

Tags: ai · tools · retries · idempotency · error-handling

You wire up a tool, ship it, and a week later a customer has two refunds for one order. The logs show one conversation, one user message, and two calls to issueRefund with identical arguments about four seconds apart.

Nobody wrote a retry loop. The retry came from somewhere else, and the tool was not built to survive being called twice.

What actually retries

Three different mechanisms can produce a second call, and they have nothing to do with each other.

1. The AI SDK's maxRetries. This retries the HTTP request to the model provider: a 429, a 500, a dropped connection. It defaults to 2, so a call can be attempted three times. Crucially, it retries the request, and the request is the whole conversation so far. If the previous step already ran your tool, the tool is not re-run by this retry; the SDK is only re-asking the model.

2. The model itself. With stopWhen: isStepCount(n), the model sees the tool result and decides what to do next. A tool that returns something the model reads as failure ({ error: "timeout" }, an empty object, a thrown error surfaced as text) is an invitation to try again. The model does not know your tool has side effects. It knows it did not get an answer.

3. Your own infrastructure. A function that times out and is retried by the platform, a user who hits send twice, a client that resubmits.

The one that surprises people is (2). It looks like a retry loop nobody wrote, because the loop is a conversation.

The wrong way

export const issueRefund = tool({
  description: "Refund an order.",
  inputSchema: z.object({ orderRef: z.string(), amount: z.number() }),
  execute: async ({ orderRef, amount }) => {
    const charge = await payments.refund({ orderRef, amount });
    await db.order.update({ where: { ref: orderRef }, data: { refundedAt: new Date() } });
    await email.send({ template: "refund", orderRef });
    return { ok: true, refundId: charge.id };
  },
});

Three failure modes are baked in:

  • If email.send throws, the refund has already happened but the tool reports failure. The model apologises and tries again. The customer gets two refunds and one email.
  • If the whole thing succeeds but the response stream drops before the model reads the result, the next attempt starts from a conversation where the tool has no result, and calls it again.
  • If the model is asked for "a refund for the last two orders" it may emit two parallel tool calls with the same orderRef because it misread the history.

The right way

Make the tool idempotent on a key the caller supplies. The AI SDK gives you toolCallId in the execute options. It is stable for one tool call, including across the SDK's own request retries.

export const issueRefund = tool({
  description:
    "Issue a refund for one order. Only call this after the customer has confirmed the amount.",
  inputSchema: z.object({
    orderRef: z.string().regex(/^ORD-[0-9]{6}$/),
    amountCents: z.number().int().positive().max(50_000),
  }),
  // AI SDK 7: the user id arrives through `toolsContext`, checked against this
  // schema, never as an argument the model fills in.
  contextSchema: z.object({ userId: z.string() }),
  execute: async ({ orderRef, amountCents }, { toolCallId, context: { userId } }) => {
    const order = await db.order.findFirst({ where: { ref: orderRef, ownerId: userId } });
    if (!order) return { ok: false, reason: "no such order on this account" };
    if (order.refundedAt) return { ok: true, alreadyRefunded: true, refundId: order.refundId };

    // The payment provider dedupes on this key: a second call with the same
    // key returns the first result instead of moving money again.
    const charge = await payments.refund({ orderRef, amountCents, idempotencyKey: toolCallId });

    await db.order.update({ where: { id: order.id }, data: { refundedAt: new Date(), refundId: charge.id } });

    // Side effects that are not part of the money movement happen after, and
    // their failure does not fail the tool.
    void email.send({ template: "refund", orderRef }).catch((error) => {
      console.error("[refund] receipt email failed", { orderRef, message: String(error) });
    });

    return { ok: true, refundId: charge.id };
  },
});

Four things changed, and each maps to one of the failure modes:

  • An idempotency key. Every payment API supports one. So does any INSERT with a unique constraint on that key. Without it, "did this already happen?" is unanswerable.
  • A pre-check that returns success. alreadyRefunded: true tells the model the goal is achieved. Returning an error here is what causes the second attempt.
  • Ordering by reversibility. The irreversible step happens once and first; the recoverable steps (email, analytics, cache invalidation) happen after and cannot fail the tool.
  • Failures are data. { ok: false, reason } ends the loop cleanly. A thrown error, by contrast, propagates as a tool error the model reads as "try differently".

Deciding what a failure should look like

SituationReturnWhy
Not found, no match, nothing to do{ ok: false, reason }The model tells the user; no retry
Already done{ ok: true, already: true }Ends the loop, no second side effect
Bad input that passed the schema{ ok: false, reason }The model can ask the user for a correction
Upstream 5xx, likely transientthrowThe SDK's retry is the right handler
Upstream 4xx, will never succeed{ ok: false, reason }Retrying wastes a step and a billed request

The rule of thumb: throw only when a retry could plausibly help. Everything else is a result.

Bound the loop

streamText({
  model: languageModel(),
  tools: { issueRefund },
  toolsContext: { issueRefund: { userId: session.user.id } },
  stopWhen: isStepCount(6),
  maxRetries: 2,
  messages,
});

isStepCount (called stepCountIs before AI SDK 7) is the wall at the end of the conversation loop; maxRetries is the transport budget for each individual request. Both are needed and they do different jobs. Six steps is a reasonable default for search-and-answer; raise it per surface for genuinely multi-step work, never globally.

Partial failure in a multi-tool step

A model can emit several tool calls in one step, and they run concurrently. If three succeed and one throws, the successful three have already happened, and there is no rollback: you are not in a transaction.

The practical mitigations, in order of how much they help:

  1. Make each tool individually idempotent (above). Then a re-run is harmless.
  2. Never expose two tools that must succeed together. If two things must be atomic, that is one tool wrapping one transaction.
  3. Return enough in each result for the model to describe what did happen. A user told "I refunded the order but could not cancel the subscription" is in a much better position than one told "something went wrong".

Checking your work

  • Call the tool twice with the same toolCallId in a test and assert the side effect happened once.
  • Force the second, non-critical side effect to throw and assert the tool still reports success.
  • Add an eval case that asks for the same action twice in one conversation and assert the second answer says it was already done.