Hello! Welcome to your next lesson in our journey through modern AI architectures.
In our last few lessons, we've focused on building and refining a powerful Retrieval-Augmented Generation (RAG) system. We started with a basic pipeline, enhanced it with a cross-encoder for reranking, and even learned how to fine-tune our own bi-encoder using contrastive learning for superior retrieval. We have built a sophisticated machine; now it's time to learn how to inspect its engine.
This lesson directly addresses that need. Our goal is to evaluate RAG systems using retrieval and generation metrics. Just as you'd write unit and integration tests in software engineering to ensure each component and the overall system works correctly, we need a robust evaluation framework to measure the performance of our RAG pipeline. This will allow us to quantify the impact of our improvements and diagnose any failures.
1. Why RAG Evaluation is Different
Evaluating a RAG system is more complex than evaluating a standard Large Language Model (LLM). A traditional LLM evaluation might focus on the quality of the generated text using metrics like perplexity or comparing it to a reference answer. However, a RAG system is a two-stage pipeline:
- Retrieval: The system first retrieves a set of documents (context) from a knowledge base.
- Generation: The LLM then synthesizes an answer based on the user's query and the retrieved context.
A failure can occur at either stage. The retriever might fail to find the relevant information, or the generator might hallucinate or misinterpret the perfectly good context it was given. Therefore, we need to evaluate both components separately before looking at the end-to-end performance.
RAG evaluation: a technical guide to measuring retrieval-augmented generation
To understand this two-part challenge, let's start with the article 'RAG evaluation: a technical guide to measuring retrieval-augmented generation' from Toloka AI. It provides a clear explanation of why RAG evaluation requires a specialized approach.
Read the first two sections, up to (but not including) 'Key components of a RAG pipeline'. Focus on the explanation of the two-stage architecture and why traditional LLM evaluation metrics fall short.
This fundamental insight leads us to a structured evaluation approach where we measure the performance of the retriever and the generator independently.
2. Evaluating the Retrieval Component
The first question to answer is: "Did we retrieve the right information?" The quality of the retrieval step is paramount; if the context is irrelevant, the generator has no chance of producing a correct answer.
Given your background in machine learning, you'll find these metrics familiar. They are adapted from the classic field of Information Retrieval.
- Context Precision: Of the documents we retrieved, what proportion are actually relevant? This measures the signal-to-noise ratio of your retrieved context. High precision means less distracting, irrelevant information is passed to the LLM.
- Context Recall: Of all the relevant documents that exist in the entire database, what proportion did we successfully retrieve? This measures how comprehensive your retrieval is. High recall means you are not missing crucial pieces of information.
These are often measured as Precision@k and Recall@k, where k is the number of documents retrieved. For example, Precision@5 measures the precision of the top 5 retrieved documents.
While precision and recall are fundamental, they don't consider the order of the retrieved documents. In RAG, we want the most relevant documents to appear at the top. For this, we use ranking-aware metrics:
- Mean Reciprocal Rank (MRR): This metric measures how high up the list the first relevant document is. It's calculated as the average of the reciprocal ranks of the first correct answer across a set of queries. An MRR of 1 means we always find a relevant document in the first position.
- Normalized Discounted Cumulative Gain (nDCG): This is a more sophisticated metric that evaluates the overall ranking quality. It assigns higher scores for retrieving highly relevant documents at top positions and penalizes for placing them lower in the list.
How to Evaluate RAG Systems: The Complete Technical ...
The article 'How to Evaluate RAG Systems' provides a concise breakdown of these core retrieval metrics.
Read the section '4.1 Retrieval Quality Metrics'. This section gives clear, practical definitions for Contextual Relevancy, Precision, Recall, Hit Rate, MRR, and NDCG.
3. Evaluating the Generation Component
Once we've confirmed our retriever is working well, we ask the next question: "Did the LLM use the retrieved context correctly?" Here, we evaluate the generated answer based on the context that was provided to it.
Two key metrics have emerged as standards for this, popularized by frameworks like RAGAS:
-
Faithfulness: This measures the factual consistency of the generated answer against the given context. It answers the question, "Is the model making things up?" or "Is the answer grounded in the provided sources?". A low faithfulness score indicates the model is hallucinating, even if the final answer sounds plausible. This is your "fake news" detector.
-
Answer Relevancy: This metric evaluates how well the generated answer addresses the original user's question. An answer can be perfectly faithful to the source material but completely miss the point of the query. This metric scores the answer's relevance to the question.
Session 7: RAG Evaluation with RAGAS and How to Improve Retrieval
The video 'Session 7: RAG Evaluation with RAGAS' from AI Makerspace provides excellent, concrete examples that illustrate the difference between faithfulness and answer relevancy.
Watch from 08:30 to 10:37. The video uses simple questions about Einstein and France to clearly demonstrate what high/low faithfulness and high/low answer relevancy look like.
Test your understanding!
Consider a RAG system for a company's internal documentation.
Query: "What is the process for requesting a new software license?"
Retrieved Context: A document detailing the company's sick leave policy.
Generated Answer: "To request sick leave, you must inform your manager and fill out form HR-123."
How would you rate the Faithfulness and Answer Relevancy of this response?
Show answer
- Faithfulness: High. The answer is factually correct and grounded in the provided context (the sick leave policy document).
- Answer Relevancy: Very Low. The answer is completely irrelevant to the user's question about software licenses.
This example clearly shows a failure in the retrieval stage, which led to a useless (but faithful) answer.
4. RAGAS: A Framework for Automated Evaluation
Measuring metrics like faithfulness and relevancy seems subjective. How can we automate this? The solution is both clever and very meta: we use a powerful LLM (like GPT-4) as an automated "judge."
The RAGAS (Retrieval-Augmented Generation Assessment) framework packages this "LLM-as-a-judge" approach into an easy-to-use library. It formalizes the four key metrics we've discussed:
- Retrieval Metrics:
context_recall: Measures if the retriever fetched all necessary information.context_precision: Measures if the retriever fetched any irrelevant information.
- Generation Metrics:
faithfulness: Measures if the answer is grounded in the context.answer_relevancy: Measures if the answer is relevant to the question.

Session 7: RAG Evaluation with RAGAS and How to Improve Retrieval
To see how these four metrics work together, let's return to the AI Makerspace video, which provides an excellent overview.
Watch the segments from 07:31 to 08:30 and 10:37 to 15:32. This will formally introduce you to context_precision and context_recall as defined by RAGAS and summarize how the four key metrics cover both the retrieval and generation aspects of the pipeline.
5. A Practical Evaluation Workflow
Now, let's outline the step-by-step process for evaluating your RAG system, as demonstrated in a typical MLOps workflow.
Step 1: Create an Evaluation Dataset
You need a "golden" dataset to test against. At a minimum, each entry in this dataset should contain:
question: A sample user query.ground_truth: The ideal, correct answer to the question.
This dataset is the bedrock of your evaluation. In many cases, it's created manually by domain experts. However, a pragmatic approach is to use an LLM to help generate a synthetic dataset of question/answer pairs from your source documents.
Step 2: Run Your RAG Pipeline
Iterate through each question in your evaluation dataset and run it through your RAG pipeline. For each question, you need to log:
answer: The final answer generated by your pipeline.contexts: The list of document chunks retrieved and passed to the generator.
Step 3: Evaluate with RAGAS
You now have all the necessary ingredients. You feed the collected data into the RAGAS evaluate function:
question(from your dataset)answer(from your pipeline)contexts(from your pipeline)ground_truth(from your dataset)
RAGAS will then use its LLM judge to compute the scores for faithfulness, answer_relevancy, context_precision, and context_recall for each item, giving you a comprehensive report on your system's performance.
This entire practical workflow is demonstrated in the AI Makerspace video. The demonstration is particularly valuable because it shows how to use this evaluation loop to compare different retrieval strategies—something we've been working on in previous lessons.
{
"intro": "This final video segment puts everything into practice. You'll see how to generate a synthetic dataset and use RAGAS to evaluate and compare three different retriever configurations.",
"resource_id": "[LINK](https://www.youtube.com/watch?v=mEv-2Xnb_Wk)",
"relevant_sections": [4, 5],
"instructions": "Watch from **16:19 to 31:52**. This is a practical walkthrough. Pay close attention to:
- The process of using GPT-4 to generate questions and ground truth answers from documents.
- The structure of the final evaluation dataset.
- How the RAGAS metrics are used to score different pipelines (basic, parent document retriever, ensemble retriever).
- How to interpret the results to decide which pipeline is best.",
"estimated_time": "16 minutes"
}
This process of creating a dataset, running the pipeline, and scoring the results is the core loop of RAG evaluation. It transforms the art of building a RAG system into a rigorous engineering discipline.
Conclusion
You now have a solid framework for measuring the quality of any RAG system. This is a critical skill that separates building a demo from deploying a production-ready AI application.
Key Takeaways:
- RAG evaluation is a two-part process, assessing both the retrieval and generation components.
- Retrieval quality is measured with metrics like Context Precision, Context Recall, MRR, and nDCG.
- Generation quality is primarily measured by Faithfulness (is the answer factual?) and Answer Relevancy (does the answer address the question?).
- Frameworks like RAGAS automate this evaluation by using a powerful LLM-as-a-judge to score these qualitative aspects.
- A robust evaluation workflow involves creating a test dataset, running your pipeline to generate outputs, and then using a framework to score the results against ground truth.
Preview of the Next Lesson:
We have now spent considerable time building, improving, and evaluating RAG systems—a powerful architecture where an LLM reasons over provided data. We are now ready to expand our horizons to systems where the LLM doesn't just reason, but acts. In the next module, we will dive into the exciting world of Agentic AI. Our first lesson will focus on designing a basic agent using a perception-action loop, the foundational concept for building LLMs that can interact with tools and external environments.