Blog / AI systems

Observability for AI features: what to log, trace and watch

When an AI feature misbehaves, you need to know which request, which prompt and which model did it, and what it cost. What to log, trace and watch so you can answer that in minutes, not days.

An AI feature can fail in ways a normal service never does. It can answer every request quickly and still be wrong, get slower and more expensive week by week, or change its behaviour because someone edited a prompt. Uptime checks see none of that. To run AI in production you need to see what it actually did, for whom, with which instructions and at what cost.

If you lead the team: what to ask

  • If a customer reports a strange answer, how long does it take us to find out what produced it?
  • Do we know what each AI feature costs per use, and would we notice if that jumped?
  • Is there a view I can read myself, showing how often people accept or correct the AI’s output?

Here is what we instrument before an AI feature goes live.

One trace per request, end to end

A single user action often becomes several steps: fetch context, build a prompt, call a model, validate the output, maybe call it again, then write a result. Give that whole chain one correlation ID and carry it through every log line, every model call and every database write. When someone reports a strange answer, you should be able to paste one ID and see the full path: what went in, what came out, how long each step took and where it went wrong.

Log inputs and outputs, with personal data removed

You cannot debug what you did not keep. Store the prompt that was actually sent (after templating) and the raw response, not just the parsed result. But AI inputs often contain names, emails, phone numbers and free text about real people, so redact before you write, not after. Decide which fields are kept, which are masked and which are never logged, and set a retention period. Logs that leak personal data turn a debugging tool into a liability.

Version on every call

Every model call should record the model and its version, the prompt version and the configuration (temperature, tools, limits) that produced it. Without this, a quality change has no explanation: was it the model provider, the prompt edit on Tuesday or a config flag? With it, you can compare before and after a change, and you can tell which outputs need to be re-checked when one version turns out to be wrong.

Latency and cost per feature, not per account

The provider bill tells you what you spent in total. It does not tell you that one feature uses most of the tokens, or that a retry loop doubled the cost of another. Track latency (median and slow tail) and token cost per feature and per step. That is what lets you choose the right model for each job and spot a runaway cost before the invoice does.

Errors, fallbacks and retries

Count the failures that do not look like failures:

  • responses rejected by schema or completeness checks;
  • retries, and how many succeed on the second attempt;
  • fallbacks taken: a smaller model, a rule-based answer, a cached result;
  • timeouts and rate-limit responses from the provider.

A rising fallback rate is often the first sign that something upstream changed, long before users complain.

Quality signals from real use

Technical metrics show the system is running. Quality signals show it is useful. The best ones come from people: a human correcting an output, rejecting it, editing it heavily or ignoring it. Record those as events tied to the trace and the prompt version. Over time they tell you where the model is reliable and where it is not, and they become the test cases for the next change.

Watch for drift

Inputs change: new products, new customer types, a new data source with different formatting. Outputs shift with them. Track simple distributions over time, such as input length, the share of each category the model assigns and average scores, and alert when they move sharply. Drift is rarely dramatic on a single day; it shows in a trend.

Alerts that matter to the business

Alert on what someone will act on: the fallback rate above a threshold, cost per feature over budget, a jump in rejected outputs, a quality signal falling. Avoid alerts that fire daily and get ignored. Each alert should have an owner and a first step written next to it.

Dashboards for people who are not engineers

Product owners and managers need a view in their own terms: how often the feature is used, how often people accept its output, what it costs per use and whether that is changing. If the only view is a log search, the people responsible for the outcome cannot see it.

Checklist before an AI feature goes live

  • One correlation ID across every step of a request.
  • Prompts and responses logged, personal data redacted, retention set.
  • Model, prompt version and config recorded on every call.
  • Latency and cost tracked per feature.
  • Rejections, retries and fallbacks counted.
  • Human corrections and rejections captured as events.
  • Drift tracked on a few simple distributions.
  • A short list of alerts, each with an owner.
  • A dashboard a non-engineer can read.

If you want to see how ready your own AI plans are for this, start with our AI readiness check.

Written from our engineers’ work on production systems. Want a second opinion on your project? Talk to an engineer.

See the work →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive, or difficult — an engineer reads every message.

Your message goes straight to our engineers at our address.