Blog / AI systems

Human in the loop, placed where it counts

Human review makes AI safe to act on, but only when it sits at the right points and people can actually do it well. Where to place reviewers, how to design the approval step, and how to turn every correction into evidence.

“A human reviews it” is the most common answer to “how do you know the AI is right?”. It is also one of the easiest to get wrong. Put a person in front of every output and they stop reading after a week. Put nobody in front of anything and the first confident mistake reaches a customer. The useful question is not whether to have a human in the loop, but where, and what that person needs to do the job well.

If you lead the team: what to ask

  • Which AI outputs go to a person before they take effect, and why those?
  • Who owns the review queue, and what happens when an item waits too long?
  • Are reviewers’ corrections saved and used to test the next change?

Where it sits on the sales-coaching product

On the AI sales-coaching product we built, the model scores sales activity against defined rules. Scores that drive an action go through a person first, and that person’s corrections are recorded. Over time, those corrections show where the model is right and where it is not: which criteria it handles well, and which still need a human eye. Review is not just a safety step; it is how the system learns where it can be trusted.

Where to put people

Review effort is limited, so spend it where a mistake costs the most:

  • Consequential actions. Anything that changes a customer’s experience, someone’s evaluation, a record of truth or money.
  • Irreversible actions. A sent email cannot be unsent. A deleted record may not come back. If it cannot be undone, a person confirms it.
  • Low confidence. When the model’s output fails a check, disagrees with a rule-based reference, or the input is incomplete, route it to a person instead of guessing.
  • Novel cases. Inputs that look unlike anything the system has handled before: a new product, a new customer type, an unusual format.

Everything else can flow without approval, as long as it is logged and can be checked afterwards.

Review queues need owners and deadlines

A review step is a workflow, and workflows need structure. Each queue should have an owner, a target time to decision and a visible backlog. Decide what happens when that time passes: escalate, hold the action or fall back to a safe default. A queue nobody owns quietly becomes the place where work stops.

Design the approval step for judgment

The interface decides whether review is real. Show the reviewer what they need to judge, not just the output:

  • the evidence the model used, quoted from the source;
  • why the item is in the queue: low confidence, a failed check, a high-impact action;
  • approve, edit and reject as equally easy options, with a short reason on rejection;
  • what will happen after approval, stated plainly.

Avoid rubber-stamping and review fatigue

When nearly everything gets approved, people stop looking. Some ways to keep review honest:

  • keep volume manageable by routing only what needs a person;
  • track approval rates and time spent per item; a reviewer approving hundreds of items in seconds each is a signal, not productivity;
  • occasionally include items with a known answer to check that review is catching errors;
  • let reviewers flag a bad rule or prompt, not just a bad output.

Corrections are your best data

Every edit or rejection is a labelled example from someone who knows the business. Store it with the original output, the input, the prompt and model version, and the reason. That gives you an evaluation set built from real cases, a way to test the next prompt change against past mistakes, and evidence for deciding when an area is reliable enough to need less review.

Human-in-the-loop checklist

  • Review required for consequential, irreversible, low-confidence and novel cases.
  • Each queue has an owner, a target time and a rule for what happens when it is missed.
  • The reviewer sees evidence and the reason for review, not only the output.
  • Approve, edit and reject are equally easy; rejections carry a reason.
  • Approval rates and time per item are tracked.
  • Corrections are stored with their context and reused as test cases.

Deciding which parts belong to software, to the model and to people is a step in how we work.

This article is drawn from a real engagement. Client details withheld; every figure comes from the client’s own data.

Read the full case study →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive, or difficult — an engineer reads every message.

Your message goes straight to our engineers at our address.