"It seems to work" is how most RAG demos are evaluated. It is also how most RAG systems fail quietly in production. A useful evaluation separates finding the right material from using it correctly, because the fixes are completely different.
Why two layers
question -> [ RETRIEVAL ] -> chunks -> [ GENERATION ] -> answer
measure: measure:
right chunks found? faithful? relevant? honest refusal?If the right chunk never reaches the model, no prompt engineering will fix the answer. If the right chunk arrives and the answer is still wrong, retrieval is not your problem.
Build a small test set first
You need questions with known answers. Fifty good ones beat a thousand generated ones.
| Field | Example |
|---|---|
| question | "Does the model's answer count toward the context window?" |
| expected source | what-is-a-context-window#what-fills-the-context-window |
| reference answer | "Yes, input and output share the same window." |
| type | factual / explanation / multi-hop / unanswerable |
Include:
- Paraphrased questions that do not reuse the document's exact words.
- Exact-term questions with codes, commands or acronyms.
- Multi-hop questions that need two sections.
- Unanswerable questions that your documents genuinely do not cover. These test whether the system admits it does not know.
Layer 1: retrieval metrics
Recall@k: for what fraction of questions does a correct chunk appear in the top k?
Mean reciprocal rank (MRR): the average of 1 / rank of the first correct chunk. It rewards putting the right chunk first, not just somewhere in the top k.
def recall_at_k(results, expected, k=5):
hits = sum(1 for q, ranked in results.items()
if any(cid in expected[q] for cid in ranked[:k]))
return hits / len(results)
def mrr(results, expected):
total = 0.0
for q, ranked in results.items():
for rank, cid in enumerate(ranked, start=1):
if cid in expected[q]:
total += 1.0 / rank
break
return total / len(results)Use these numbers to choose chunk size, k, and whether hybrid search helps. Change one thing at a time.
Layer 2: answer metrics
For generated answers, three questions matter most:
- Faithfulness (groundedness): is every claim in the answer supported by the retrieved sources? An answer can be true but unfaithful if it came from the model's memory instead of your sources.
- Answer relevance: does it actually answer the question asked?
- Correct refusal: for unanswerable questions, does it say the sources do not cover it, instead of inventing something?
Citation checks are a cheap extra: does each cited source exist in what was retrieved, and does it support the sentence it is attached to?
Grading answers
- By hand: slow but trustworthy. Always do some.
- With an LLM as judge: give a strong model the question, the sources and the answer, and ask it to label unsupported claims. Open-source frameworks such as RAGAS package metrics like these.
Treat an LLM judge like any instrument: check it against human grades on a sample, and keep its prompt and model fixed while comparing system versions. Judges can prefer longer answers or their own writing style.
A simple evaluation loop
- Freeze a test set.
- Run the full pipeline and save retrieved chunk IDs, answers and token counts.
- Compute recall@k, MRR, faithfulness and refusal accuracy.
- Change one component.
- Re-run and compare. Keep the change only if it helps without hurting another metric.
Record cost and latency too. A setup that improves faithfulness by one point but doubles cost may not be worth it.
Summary
- Evaluate retrieval and generation separately; they fail for different reasons.
- Start with a small, hand-checked test set that includes unanswerable questions.
- Use recall@k and MRR for retrieval; faithfulness, relevance and refusal for answers.
- Treat LLM judges as instruments that also need validation.
This mindset, measuring on data the system has not been tuned on, is the same one behind train, validation and test sets in machine learning.