LLM evaluation
RAG Pipeline Evaluation: Retrieval, Answers, and Citation Quality
A reproducible checklist for retrieval-augmented generation, from evidence recall to grounded answers and failure review.
A retrieval-augmented answer can fail because the right passage was never found, because the model ignored it, or because the source document was parsed incorrectly. Measure those stages separately.
Create answerable and unanswerable questions
Sample real tasks across document formats and topics. For each question, label the passages that actually support an answer and what a correct answer must include. Add questions whose answer is absent; they test whether the system admits uncertainty. Keep this set separate from prompt and retrieval tuning.
Measure retrieval first
At a stated top-k, report whether supporting evidence appears in retrieved chunks and how much irrelevant text arrives with it. Inspect failures caused by PDF parsing, chunk boundaries, metadata filters, or stale indexes. Ragas documents context recall and context precision as retrieval measures; use human-checked labels for consequential decisions.
Then evaluate the answer
Check factual correctness against the source, whether each material claim has a matching citation, and whether the system refuses to invent an answer when evidence is missing. Faithfulness to retrieved context is different from truth if the source itself is outdated. Record model and index versions, prompt, top-k, latency, and cost. A useful report includes failure examples and a small list of fixes tied to their stage in the pipeline.