Blog / AI systems

The cost of an AI feature: tokens, caching and the right model per task

What an AI feature really costs is a number per unit of value: per lesson, per answer, per document. How to estimate it, get the quality right first, then bring the cost down with model routing, caching and batching.

“How much will the AI cost?” is usually asked as a monthly bill. The more useful question is narrower: what does one unit of value cost? One lesson written, one question answered, one document processed. Once you know that number, you can decide whether the feature pays for itself, and you can see what moves it.

If you lead the team: what to ask

  • What does one unit of value cost us: one answer, one document, one lesson?
  • Did we choose the model for the quality we need first, and only then look at price?
  • Is there an alert that catches a sudden jump in an AI feature’s cost before the invoice does?

Estimate per unit, not per month

Models charge by tokens, the pieces of text they read and write, and output tokens usually cost more than input tokens. For one unit of value, add up:

  • Input: the instructions, the context you send (documents, history, examples) and the request itself.
  • Output: what the model writes back, including anything you throw away.
  • Calls: how many model calls one unit takes. A feature that plans, writes and then checks makes three calls, not one.
  • Retries: the share of calls that fail validation and run again.

Apply the model’s list price and you have a cost per unit. Multiply by the volume you expect and you have a budget. Measure it on a sample of real inputs rather than a guess: context size varies far more than people expect.

Quality first, then cost

The cheapest model is tempting, and it is the wrong place to start. A cheap answer that is wrong costs more than an expensive one that is right, because someone has to catch it, fix it or live with it.

SuperEscuela, a free learning platform for children, is a good example. Its lessons were first generated with the smallest model. After reviewing the output against what children need (correct, clear, pitched at their level), generation moved to a higher-quality model. With that model, a lesson costs about 2–4 cents in AI: a whole school’s lessons for roughly $60–130. The saving from the smaller model was never worth the quality it gave up.

So the order is: define what a good output looks like, find the model that reaches it, and only then work on the price.

Route each task to the right model size

One feature is often several tasks of very different difficulty. Classifying a request, extracting fields or checking a format are jobs a small, fast model often handles well. Planning, writing for a demanding audience or reasoning over long documents may need a larger one. And some steps don’t need a model at all: SuperEscuela’s content pipeline uses AI in only two of its four stages, and the other two are plain code, which costs nothing per run and behaves the same every time.

Test each task on its own, and give each one the smallest model that meets its bar.

Stop paying twice for the same thing

  • Cache repeated context. If every call sends the same long instructions, examples or reference document, put them at the start and use the provider’s prompt caching. Cached input is billed at a fraction of the normal price.
  • Batch what isn’t urgent. Content generation, nightly enrichment and back-fills can go through batch APIs, which providers offer at a discount in exchange for slower delivery.
  • Trim the context. Send the three relevant paragraphs, not the whole manual; the last few turns, not the whole conversation. Smaller context is cheaper and often more accurate.
  • Constrain the output. Ask for structured output with only the fields you need. Long, chatty answers are paid for in the most expensive tokens.

Track it like any other cost

Log tokens and cost on every call, tagged by feature and by unit: this lesson, this ticket, this customer. Then you can answer what each feature costs, which one is growing, and whether a prompt change made things cheaper or more expensive.

Set budget alerts per feature and per day, so a loop that retries forever or a prompt that suddenly doubles in size is caught within hours, not on the invoice. And review the numbers when a new model comes out: the right model for a task today may not be the right one in six months, in either direction.

The short version

Price the unit of value. Reach the quality bar first. Then route, cache, batch and trim, and watch the cost per feature every day.

See how this played out across a whole curriculum in the SuperEscuela case study.

This article is drawn from a real engagement. Client details withheld; every figure comes from the client’s own data.

Read the full case study →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive, or difficult — an engineer reads every message.

Your message goes straight to our engineers at our address.