Blog / AI systems

How to test a RAG system before your team trusts it

A RAG demo answers the questions its builders thought to ask. How to test retrieval and answers separately, check citations automatically and agree a pass mark in advance, so your team can rely on what it says.

A RAG system that looks good in a demo has answered the questions its builders thought to ask. Your team will ask different ones, and they will stop using it the first time it answers confidently from the wrong document. Trust is earned before launch, with a test you can run again and again.

If you lead the team: what to ask

  • Was it tested on real questions from our people, with answers checked by someone who knows the subject?
  • What counts as good enough, and did we agree on it before testing?
  • When the prompt, the model or the documents change, is the full test run again?

Start with real questions and known answers

Collect the questions your people actually ask: from support tickets, internal chat, and the experts who answer them today. For each one, write down the correct answer and the source documents that contain it. Include hard cases on purpose:

  • answers spread across two documents;
  • questions that hinge on an exact reference number or name;
  • topics where two documents disagree;
  • questions the knowledge base can’t answer, where the right response is “I don’t know”.

Have someone who knows the subject sign off on the answers. This set is your known answer, and it is worth more than any generic benchmark because it is made of your questions and your documents.

Measure retrieval and answers separately

When an answer is wrong, you need to know which half failed.

  • Retrieval: for each question, did the system fetch the source documents you listed, and how high did they rank? If retrieval misses, no prompt will fix the answer.
  • Answer: given what was retrieved, is the answer correct, is it grounded (every claim supported by the retrieved passages), and is it cited? A correct answer that isn’t supported by its sources is an accident, and it won’t repeat.

Scoring the two separately tells you where to work: chunking and search, or prompt and model.

Check citations automatically

Every cited passage should exist, belong to a document the user is allowed to see, and actually contain the claim it is attached to. The first two are plain code checks. The third can be checked by matching quoted text against the source, or by a second model, with a person spot-checking the results. Run these on every answer in the test set. A citation that points nowhere is a defect, not a style issue.

Run the whole set on every change

A new prompt, a new model version, a re-chunked index or a fresh batch of documents can each improve some answers and quietly break others. Run the full test set as a regression on every one of those changes, and compare with the last accepted run. Look at what got worse, not just the average. Keep the run history, so you can tell when a question started failing and what changed that day.

Sample real answers for human review

Automated checks catch what you thought to check. Once the system is in use, have someone who knows the subject review a regular sample of real answers, with their sources. Each mistake they find becomes a new question in the test set, so the same failure can’t come back unnoticed.

Agree the pass mark before you test

Decide with the people who will rely on the system what “good enough” means, and write it down before the first run: how often retrieval must find the right source, how often answers must be correct and cited, and which questions must never be answered wrong at all, such as anything touching money, health or legal deadlines. Setting the bar after seeing the results turns a test into a negotiation.

The same discipline on a product we built

On the AI sales-coaching product we built, every score shows the evidence quoted from the call, so a manager can check it in seconds. Outputs are structured and validated before they are stored, and a rule-based scorer runs alongside the AI as a deterministic reference. When we tested the product’s metrics against a known answer, the test caught errors before any manager relied on them. That is the habit to bring to RAG: evidence on every output, and a known answer to check it against.

RAG test checklist

  • Real questions, with expert-approved answers and source documents
  • Hard cases, including questions with no answer
  • Retrieval scored separately from answers
  • Answers checked for correctness, grounding and citations
  • Automatic citation checks on every answer
  • A full regression run on every prompt, model or index change
  • Regular human review of real answers, feeding the test set
  • A pass mark agreed in writing before testing

Testing like this is built into how we deliver AI on your data.

This article is drawn from a real engagement. Client details withheld; every figure comes from the client’s own data.

Read the full case study →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive, or difficult — an engineer reads every message.

Your message goes straight to our engineers at our address.