Blog / AI systems

The AI harness: the code around the model is what makes it reliable

A model call is easy to add and hard to depend on. The harness around it, from prompt templates and schema validation to fallbacks and provenance, is what turns a model into a feature people can act on.

Adding a model call to a product takes an afternoon. Making that call dependable enough that people act on its answers takes engineering. Most of that engineering is not in the model at all. It is in the code around it: the harness.

If you lead the team: what to ask

  • When the AI provider is slow or down, what do users see, and is it clearly labelled?
  • Is every AI answer checked before anyone uses it, or shown as it comes?
  • If someone asks why the system gave a particular answer, can we trace what produced it?

The idea that helps most is a simple reframing: the model is a component; the harness is the product. Models get replaced, upgraded and repriced. The harness is what your users actually experience, and what decides whether a bad answer reaches them.

What the harness is made of

  • Prompt templates. Prompts live in version control, with named variables, not as strings scattered through the code. A change to a prompt is reviewed like a change to code, because it is one.
  • Structured outputs and schema validation. Ask the model for fields, not paragraphs, and check every response against a schema before anything uses it. A missing field or a value out of range is caught in code, not discovered by a user.
  • Tool calling. When the model needs data or needs to act, it requests a defined tool with typed arguments. Your code runs the tool, checks permissions and returns the result. The model asks; the harness decides.
  • Retries and timeouts. Model providers have slow moments and outages. Every call gets a timeout, a bounded number of retries with backoff, and a clear outcome when they run out.
  • Fallbacks. Decide in advance what happens when the model cannot answer in time: a simpler method, a cached result, or an honest “not available yet”. Never a silent blank, and never a guess dressed as an answer.
  • Logging and provenance. Record what went in and what came out: the inputs, the prompt version, the model, the result. When someone asks why the system said something, the answer exists.
  • Tests. A set of known inputs with expected outputs, run whenever the prompt, the model or the code changes, so a regression is found before release.

A harness in practice

On the AI sales-coaching product we built, the model scores sales calls, and those scores reach people who make decisions with them. The harness around it follows the list above:

  • The model returns structured fields: a score, the reasons for it, and evidence quoted from the call. A quote can be checked against the call itself; a vague impression cannot.
  • The output is validated before it is stored or shown.
  • A rule-based scorer runs alongside the AI. When the model is slow or unavailable, the rule-based score takes over, and it is labelled as a fallback, so nobody mistakes it for the full AI assessment.
  • Every score records its inputs, so any result can be traced back to what produced it.

None of these pieces is exotic. Together they mean an outage at a model provider degrades the product instead of breaking it, and a strange score can be investigated instead of argued about.

Why the harness outlasts the model

A good harness makes model changes routine. With schemas, tests and provenance in place, trying a newer or cheaper model is an experiment with a clear pass or fail, not a leap of faith. Without them, every model change is a risk nobody can measure, so teams either stop improving or ship changes blind.

It also makes the system explainable to the business. Managers do not need to understand the model; they need to know that its answers are checked, that there is a defined behaviour when it fails, and that any result can be traced. That is a property of the harness, and it is the part we put the most care into when we build AI on a company’s own data.

Harness checklist

  • Are prompts versioned and reviewed like code?
  • Does the model return structured fields, validated against a schema?
  • Are tool calls typed, permission-checked and run by your code?
  • Does every call have a timeout and bounded retries?
  • Is there a defined fallback, clearly labelled to users?
  • Does every result record its inputs, prompt version and model?
  • Do regression tests run when the prompt, model or code changes?

This article is drawn from a real engagement. Client details withheld; every figure comes from the client’s own data.

Read the full case study →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive, or difficult — an engineer reads every message.

Your message goes straight to our engineers at our address.