Packet To Sniff

How to evaluate a RAG system: retrieval and answer quality

Measure a RAG system in two layers: did retrieval find the right passages, and is the answer faithful to them? Learn recall@k, MRR, faithfulness and a practical test set.

By Packet To SniffPublished 3 min read

Helpful background: What is retrieval-augmented generation (RAG)?, How vector search works: similarity, top-k and hybrid search.

On this page

"It seems to work" is how most RAG demos are evaluated. It is also how most RAG systems fail quietly in production. A useful evaluation separates finding the right material from using it correctly, because the fixes are completely different.

Why two layers

Output
question -> [ RETRIEVAL ] -> chunks -> [ GENERATION ] -> answer
               measure:                   measure:
               right chunks found?        faithful? relevant? honest refusal?

If the right chunk never reaches the model, no prompt engineering will fix the answer. If the right chunk arrives and the answer is still wrong, retrieval is not your problem.

Build a small test set first

You need questions with known answers. Fifty good ones beat a thousand generated ones.

Field Example
question "Does the model's answer count toward the context window?"
expected source what-is-a-context-window#what-fills-the-context-window
reference answer "Yes, input and output share the same window."
type factual / explanation / multi-hop / unanswerable

Include:

  • Paraphrased questions that do not reuse the document's exact words.
  • Exact-term questions with codes, commands or acronyms.
  • Multi-hop questions that need two sections.
  • Unanswerable questions that your documents genuinely do not cover. These test whether the system admits it does not know.

Layer 1: retrieval metrics

Recall@k: for what fraction of questions does a correct chunk appear in the top k?

Mean reciprocal rank (MRR): the average of 1 / rank of the first correct chunk. It rewards putting the right chunk first, not just somewhere in the top k.

Python
def recall_at_k(results, expected, k=5):
    hits = sum(1 for q, ranked in results.items()
               if any(cid in expected[q] for cid in ranked[:k]))
    return hits / len(results)
 
def mrr(results, expected):
    total = 0.0
    for q, ranked in results.items():
        for rank, cid in enumerate(ranked, start=1):
            if cid in expected[q]:
                total += 1.0 / rank
                break
    return total / len(results)

Use these numbers to choose chunk size, k, and whether hybrid search helps. Change one thing at a time.

Layer 2: answer metrics

For generated answers, three questions matter most:

  1. Faithfulness (groundedness): is every claim in the answer supported by the retrieved sources? An answer can be true but unfaithful if it came from the model's memory instead of your sources.
  2. Answer relevance: does it actually answer the question asked?
  3. Correct refusal: for unanswerable questions, does it say the sources do not cover it, instead of inventing something?

Citation checks are a cheap extra: does each cited source exist in what was retrieved, and does it support the sentence it is attached to?

Grading answers

  • By hand: slow but trustworthy. Always do some.
  • With an LLM as judge: give a strong model the question, the sources and the answer, and ask it to label unsupported claims. Open-source frameworks such as RAGAS package metrics like these.

Treat an LLM judge like any instrument: check it against human grades on a sample, and keep its prompt and model fixed while comparing system versions. Judges can prefer longer answers or their own writing style.

A simple evaluation loop

  1. Freeze a test set.
  2. Run the full pipeline and save retrieved chunk IDs, answers and token counts.
  3. Compute recall@k, MRR, faithfulness and refusal accuracy.
  4. Change one component.
  5. Re-run and compare. Keep the change only if it helps without hurting another metric.

Record cost and latency too. A setup that improves faithfulness by one point but doubles cost may not be worth it.

Summary

  • Evaluate retrieval and generation separately; they fail for different reasons.
  • Start with a small, hand-checked test set that includes unanswerable questions.
  • Use recall@k and MRR for retrieval; faithfulness, relevance and refusal for answers.
  • Treat LLM judges as instruments that also need validation.

This mindset, measuring on data the system has not been tuned on, is the same one behind train, validation and test sets in machine learning.

Frequently asked questions

What is recall@k in RAG?

Recall@k is the fraction of test questions for which at least one correct source chunk appears in the top k retrieved results. If recall@5 is 0.9, the right material reached the top five for 90 percent of questions.

Can I use an LLM to grade RAG answers?

Yes, and it scales well, but treat the grader as a measuring instrument that also needs checking. Grade a sample by hand, compare with the model's grades, and keep the grading prompt and model fixed between experiments.

Continue learning

Tags

  • How to chunk documents for RAG

    Chunking decides what your RAG system can retrieve. Compare fixed-size, recursive, heading-aware and semantic chunking, and learn how to pick size and overlap.

    Retrieval-augmented generationIntermediate3 min
  • Data leakage in machine learning: how good scores lie

    Data leakage lets outside information slip into training, inflating scores that collapse in production. Learn the common types and how to prevent them.

    Machine learning methodsIntermediate3 min
  • Precision, recall and F1 score explained with an example

    Accuracy can be misleading on imbalanced data. Learn the confusion matrix, precision, recall, F1 and thresholds through a worked phishing-detection example.

    Machine learning methodsBeginner3 min