Blog / AI systems

RAG that answers from your data, not around it

Most RAG failures are data failures that retrieval exposes. What must be true before retrieval, from sources of truth and permissions to citations and hybrid search, so your assistant answers from your knowledge base.

Retrieval-augmented generation, or RAG, promises an assistant that answers from your own documents instead of from whatever the model picked up in training. The demo is easy: load some files into a vector database, ask a question, get a fluent answer. The trouble starts when the answer is fluent and wrong, or right but drawn from a document that user should never have seen.

If you lead the team: what to ask

  • For each kind of information, which document is the official one the assistant answers from?
  • Can the assistant show someone a document they couldn’t open themselves?
  • Does every answer show its sources, and does it say “I don’t know” when they aren’t there?

Most of those failures are not retrieval failures. They are data failures that retrieval exposes. What is true before retrieval decides whether a RAG system answers from your data or around it.

Our part, stated exactly

On an AI platform for legal and healthcare practitioners, we owned the integration and data layer: ingestion from many systems, normalization into one validated model, reliability and traceability, with regulated data handled to HIPAA, SOC 2 and GDPR expectations. The platform team owned search, retrieval and the agents. Working beneath them showed us how much of retrieval quality is decided upstream.

1. Name your sources of truth, and their freshness

If the same policy exists as a PDF, a wiki page and an old email, which one does the assistant quote? Decide, for each type of information, which system is the source of truth, and index that. Then decide how fresh it must be: a price list that changes daily needs a different sync from a handbook that changes yearly. Keep the last-updated time on every record so an answer can say how current it is.

2. Normalize, with stable identifiers

A knowledge base fed from several systems will hold the same customer, case or product under different IDs and spellings. Normalize them into one model with stable identifiers before indexing. Otherwise retrieval returns three half-answers about what it believes are three different things, and the model stitches them into one confident mistake.

3. Chunk with metadata, not just by length

Chunking splits documents into pieces small enough to retrieve. Cutting at a fixed length splits tables in half and separates a heading from the paragraph that explains it. Follow the document’s own structure instead: sections, clauses, notes. Then attach metadata to every chunk: source system, document ID, title, section, date and access rules. Embeddings find text that sounds similar; metadata is what lets you filter it, cite it and secure it.

4. Enforce document-level permissions per user

An assistant must never retrieve what the person asking couldn’t open themselves. Carry each document’s permissions into the index and filter on them at query time, for that user, before anything reaches the model. Filtering after generation is too late: the model has already read it. With regulated data, this is not a feature. It is the requirement.

5. Use hybrid search

Vector search over embeddings is good at meaning: it matches “cancel my plan” to “terminate the subscription”. It is weak at exact things: case numbers, product codes, names, rare terms. Keyword search is the opposite. Hybrid search runs both and merges the results, so a question about a specific reference number finds that record, and a loosely worded question still finds the right section.

6. Cite every answer

Every claim in an answer should point to the chunk it came from, with a link to the source document. Citations let a user verify an answer in seconds, and they let you test automatically whether the answer is supported by what was retrieved. An answer without a source is an opinion.

7. Answer from sources, or say you don’t know

Instruct the model to answer only from the retrieved passages, and to say so plainly when they don’t contain the answer. Then enforce it in code: if the retrieved context is empty or weak, return “I couldn’t find this in your documents” instead of letting the model fill the gap from memory. A clear “I don’t know” builds more trust than a plausible guess.

Before you index anything

  • A named source of truth for each type of information
  • A freshness target, and a last-updated time on every record
  • One normalized model with stable identifiers
  • Structure-aware chunks with source, date and access metadata
  • Permission filtering per user, at query time
  • Hybrid keyword and vector search
  • Citations on every answer
  • A tested “I don’t know” path

If you want an assistant that answers from your own data, start with the AI readiness check, or see how we deliver AI on your data.

This article is drawn from a real engagement. Client details withheld; every figure comes from the client’s own data.

Read the full case study →

Want us to look
at your site?

Tell us where traffic, revenue or your numbers stopped making sense. We will tell you what we would check first.

Prefer to write directly? enable JavaScript to see the address

Talk to an engineer

No sales theater. Tell us where your operation feels slow, repetitive, or difficult — an engineer reads every message.

Your message goes straight to our engineers at our address.