Skip to content

When LLM tracing pays for itself, and the free version to build first

A user reports an answer you cannot reproduce. Without the exact prompt, the tool steps and the model version, you are guessing. But a paid tracing tool is not the first thing to reach for.

AI bundle5 min readships at docs/solutions/ai-bundle/when-tracing-pays-for-itself.md

Tags: ai · observability · tracing · logging · debugging

A user sends a screenshot: the assistant told them their subscription was cancelled. It was not. You try the same question and get a correct answer. You try it five more times. Correct every time.

You now have no way to find out what happened, because the only record of that conversation is a screenshot. Not the messages that were in context, not the tool calls, not which model version served it, not the system prompt as it was before yesterday's edit.

That gap is the case for tracing. The question is when it justifies a dependency and a recurring bill.

The three signals

1. You cannot reproduce a reported answer. This is the one that matters. Non-reproducible reports mean the inputs are not the same as the ones you have, and you cannot see the inputs. If this has happened twice, you are past due.

2. You have more than one prompt in production and you are changing them. With one prompt and no history you can reason about it. With five prompts, two models and a tool loop, "which version answered this?" needs a record.

3. Quality is a metric someone tracks. The moment anyone asks "is it getting better?", you need per-call outcomes joined to the prompt version. That is what tracing tools are actually for; debugging is a side effect.

Conversely, do not add it on day one. A prototype with one prompt and ten users gets more from an eval suite than from a trace viewer.

Build the free version first

Most of the value is structured logging you can add in an hour, with no vendor. The thing that makes it useful is a conversation id that appears on every record and is surfaced in the UI, so a user can quote it.

const conversationId = parsed.value.conversationId ?? crypto.randomUUID();

const result = streamText({
  model: languageModel(parsed.value.modelId),
  instructions: systemPrompt(),
  messages: await toModelMessages(parsed.value.messages),
  abortSignal: request.signal,
  // AI SDK 7 names: `onEnd` (was `onFinish`); `usage` covers every step and
  // the per-request metadata lives on `finalStep`.
  onEnd({ usage, finishReason, finalStep }) {
    after(
      logAiCall({
        conversationId,
        feature: "chat",
        // The version of the prompt, not the prompt. See below.
        promptVersion: PROMPT_VERSION,
        modelId: finalStep.response.modelId,
        // The provider's own id for the request. Support can look it up.
        providerRequestId: finalStep.response.id,
        inputTokens: usage.inputTokens ?? 0,
        outputTokens: usage.outputTokens ?? 0,
        finishReason,
        durationMs: Date.now() - startedAt,
      }),
    );
  },
  onError({ error }) {
    console.error("[chat] stream error", { conversationId, error });
  },
});

That is enough to answer most questions: which model, which prompt version, how long, how many tokens, did it finish or hit the length ceiling, and the provider's request id for a support ticket.

Version your prompts. A constant next to the prompt, bumped when the text changes, costs nothing and turns "quality dropped last Tuesday" into a join.

What not to log

This is where a debugging feature becomes a data-protection incident.

  • Never log message content by default. Chat messages are whatever the user typed: names, addresses, medical questions, card numbers.
  • Never log tool arguments. They are derived from user text.
  • Never log the raw failing output from a structured-generation error. It contains what the user pasted.
  • Never log the API key, obviously, but also never log the full request headers, which is how keys end up in logs by accident.

If you need content to debug, make it opt-in and bounded: a per-user "share this conversation with support" action that stores it with an expiry, or a sampled 1% with PII scrubbing applied on the way in. "We will redact it later" is not a plan; logs are copied to places you do not control.

When to buy the tool

Reach for a hosted LLM observability tool (Langfuse, Braintrust, LangSmith, Helicone, or your existing APM's LLM module) when at least two are true:

  • You need per-step traces of tool loops. Reconstructing a six-step agent run from log lines is genuinely painful, and this is where the tools earn their price.
  • You want prompt versions and evals in the same place, so a score can be attributed to a prompt change.
  • Someone non-engineering needs to read the traces: support, or whoever owns quality.
  • You are ready to annotate traces with outcomes (thumbs up, resolved, escalated). Traces without outcomes are just expensive logs.

In AI SDK 7 OpenTelemetry lives in its own package. Register it once (in Next.js, in instrumentation.ts) and every call emits spans; telemetry: { functionId: "chat" } on a call names them:

import { OpenTelemetry } from "@ai-sdk/otel";
import { registerTelemetry } from "ai";

registerTelemetry(new OpenTelemetry());

Add the vendor's exporter next to it. That is an hour of work. The cost is not the integration; it is the recurring bill, a vendor holding your prompts and possibly your users' text, and one more thing to keep working.

Cost, honestly

These tools price per trace or per span, and an agent turn is many spans. A chat feature doing 100k turns a month will not be on the free tier. Two things keep it affordable:

  • Sample. Trace 100% of errors and 1-5% of successes. You do not need every happy path.
  • Do not trace what you do not read. A dashboard nobody opens is a subscription, not observability.

The minimum that makes a report actionable

Whatever you build, a user report should let you find:

  1. the conversation id (shown in the UI, quotable in a bug report)
  2. the prompt version and model id that served it
  3. the tool calls made, by name, with outcomes
  4. finishReason: "length" explains a surprising number of "it stopped making sense" reports
  5. the provider request id, for escalating to the provider

None of that requires content logging, and none of it requires a vendor. Add the vendor when the tool-loop traces or the shared workspace are what you are missing, not before.