Skip to content

How to Evaluate a RAG System: Retrieval Quality, Groundedness and Hallucination

How to Evaluate a RAG System: Retrieval Quality, Groundedness and Hallucination

How to Evaluate a RAG System: Retrieval Quality, Groundedness and Hallucination

RAG systems are easy to demo and much harder to trust in production. A polished answer can still be wrong because the system retrieved the wrong evidence, missed the right document, supplied too much context or generated a claim that was not supported by the source material.

That is why retrieval-augmented generation needs to be evaluated as a complete system rather than judged by whether a few example answers look good.

What should a RAG system be evaluated on?

A useful evaluation framework separates retrieval quality from generation quality. If those two layers are measured together, it becomes difficult to diagnose why an answer failed.

1. Retrieval quality

The first question is whether the system found the right evidence. Useful measures include whether the correct document or passage appears in the retrieved results, how highly it ranks, and whether irrelevant material is being pulled into the context window.

For document-heavy applications, this is often where the biggest gains come from. Chunking, metadata, hybrid search, filtering, re-ranking and query transformation can all change retrieval quality before the language model generates a single word.

2. Groundedness

A grounded answer should be supported by the evidence the system retrieved. The response should not introduce claims that cannot be traced back to the source material.

This matters especially in financial services, healthcare, legal, compliance and other regulated or policy-heavy environments where users need to understand why the system reached an answer.

3. Hallucination rate

Hallucination is not just a model problem. A model may invent information because it received weak evidence, conflicting passages or incomplete context. Evaluation should therefore record whether unsupported claims came from generation itself or from poor retrieval upstream.

4. Source traceability

Production RAG should preserve where an answer came from. Depending on the application, that may include the source document, section, passage, URL, document version or citation presented to the user.

Traceability gives users a way to verify important answers instead of treating fluent model output as authority.

5. Consistency

A useful system should answer the same underlying question consistently even when users phrase it differently. Testing paraphrases, follow-up questions and ambiguous prompts helps expose brittle retrieval behaviour that a small demo set can miss.

6. Permission correctness

In enterprise RAG, an accurate answer can still be a serious failure if the user was not authorised to see the retrieved information. Evaluation should include access-control tests alongside answer-quality tests.

7. Latency and cost

Accuracy is only part of production quality. Retrieval depth, re-ranking, model choice and context size all affect response time and inference cost. A production system needs to find the smallest amount of high-quality context required to answer reliably.

Build a representative evaluation set

The most useful evaluation data comes from the questions real users are expected to ask. A good test set should include straightforward questions, edge cases, ambiguous requests, questions that require several pieces of evidence, and questions the system should refuse to answer because the knowledge base does not contain enough information.

For each question, teams can record the expected source evidence and the characteristics of an acceptable answer. This creates a repeatable benchmark that can be rerun when chunking, embeddings, ranking, prompts or models change.

Evaluate retrieval and generation separately

When a RAG answer is wrong, start by asking whether the correct evidence was retrieved.

If the evidence was missing, the problem is likely in ingestion, chunking, indexing, metadata, filtering or ranking. If the right evidence was present but the answer was still wrong, the problem is more likely in context construction, prompting or generation.

This separation makes optimisation much more systematic than repeatedly changing prompts and hoping the output improves.

RAG evaluation in regulated workflows

In regulated environments, evaluation should go beyond answer quality. Teams should also test source traceability, permissions, auditability, versioning and how the system behaves when guidance conflicts or changes.

Our work on Vellum, a mortgage underwriting platform, involved retrieval and document intelligence across more than 3,000 pages of mortgage guidance. In that kind of workflow, the usefulness of the AI depends heavily on whether the right rule, condition or exception is retrieved and can be checked by the user.

Read our production RAG engineering deep dive for more on ingestion, chunking, retrieval and production architecture.

A practical production RAG evaluation loop

  1. Collect representative user questions.
  2. Define the expected evidence for each question.
  3. Measure whether retrieval finds and ranks that evidence correctly.
  4. Evaluate whether the generated answer stays grounded in the retrieved context.
  5. Record unsupported claims, missed evidence and permission failures.
  6. Track latency and model cost alongside quality.
  7. Use real user feedback and failed queries to expand the evaluation set over time.

RAG evaluation FAQs

How do you measure RAG accuracy?

Measure retrieval and generation separately. First check whether the correct evidence was retrieved and ranked highly enough. Then check whether the final answer is supported by that evidence and avoids unsupported claims.

What is groundedness in RAG?

Groundedness describes whether the generated answer is supported by the retrieved source material. A grounded response should be traceable back to evidence the system actually retrieved.

How can RAG hallucinations be reduced?

Improve document quality, retrieval, ranking and context selection, then evaluate the model against representative questions. Hallucinations often start with weak or incomplete retrieval rather than the language model alone.

Do RAG systems need continuous evaluation?

Yes. Documents, models, prompts and user behaviour change over time, so a production RAG system should be evaluated continuously rather than only before launch.

Building a RAG system that can be trusted

If you are designing a RAG product for complex documents, internal knowledge or regulated workflows, see our RAG development services and our UK RAG development cost guide.