Chatbot & RAG Evaluation
- Eyeballing doesn't scale — you need automated metrics to catch regressions and compare changes.
- Four key metrics: faithfulness (grounded in docs?), answer relevancy (answers the question?), context precision (right docs ranked high?), context recall (all needed docs retrieved?).
- RAGAS and DeepEval compute these metrics using LLM-as-judge. Build a golden dataset of Q&A pairs, run eval in CI, and track scores over time.
You changed the chunking strategy. Or swapped the embedding model. Or tweaked the prompt. Did things get better or worse? Without evaluation, you're guessing. And with non-deterministic outputs, your gut feeling is wrong more often than you think. Evaluation gives you numbers to make decisions on. One submodule per idea, ending with a cheat sheet.
Why eval matters
RAG systems are non-deterministic — the same question can produce different answers across runs. This creates three problems:
- Regressions are invisible — you change one thing, and something unrelated breaks. Without eval, you won't notice until a user complains.
- A/B decisions are gut-feel — "GPT-4o-mini vs Claude Haiku for our RAG pipeline?" Without metrics, you're comparing vibes.
- Production drift — the system works today. In three months, with new documents and new query patterns, it quietly degrades.
Automated evaluation catches all three.
The four key metrics
These come from the RAG evaluation literature and are implemented by frameworks like RAGAS.
Faithfulness
Does the answer stick to what the retrieved documents say?
A faithfulness score of 1.0 means every claim in the answer is supported by the context. A low score means the model is hallucinating — making up information that isn't in the docs.
How it's measured: The evaluator LLM extracts individual claims from the answer, then checks each one against the retrieved context. Faithfulness = (supported claims) / (total claims).