Hello! Welcome to the next lesson in our journey through Retrieval-Augmented Generation (RAG).
In our previous lesson, we successfully built a complete, standard RAG pipeline. We used a dense retriever (a bi-encoder) to fetch relevant documents and an LLM to generate an answer based on them. We also touched upon the limitations of this "naive" approach and previewed that the RAG landscape is rich with more advanced, robust techniques.
A key limitation of our initial retriever is that while it's fast, it prioritizes "semantic similarity," which isn't always the same as "relevance." A document might use similar words to your query but fail to answer the actual question. This can lead to the generator receiving suboptimal context.
This lesson directly addresses that problem. Our goal is to implement a cross-encoder model for reranking retrieved documents. We will introduce a second stage to our retrieval process—a high-precision reranking step—to ensure the documents passed to our generator are not just similar, but truly relevant.
1. The Two-Stage Retrieval Paradigm: Recall and Precision
In information retrieval, there's often a trade-off between recall (finding all relevant documents) and precision (ensuring the found documents are relevant).
- Stage 1: Fast Retrieval (High Recall): This is what our bi-encoder from the last lesson does. It quickly scans millions of documents to retrieve a broad set of candidates (e.g., the top 25-50 most similar documents). Its strength is speed and its ability to "cast a wide net."
- Stage 2: Reranking (High Precision): This stage takes the smaller candidate set from Stage 1 and applies a more computationally intensive but far more accurate model to re-order them based on true relevance.
This two-stage process allows us to get the best of both worlds: the speed and scalability of a bi-encoder and the accuracy of a more powerful model. The model we'll use for this second stage is a cross-encoder.
To start, let's watch a short video that introduces the concept of reranking and why it's a valuable addition to a RAG pipeline.
Do Reranking Models Actually Improve RAG?
The video 'Do Reranking Models Actually Improve RAG?' by Adam Lucek provides a great high-level overview of reranking and introduces the cross-encoder architecture.
Watch the first two minutes of this video (from 00:00 to 02:00). Pay attention to why a second reranking stage is needed and how it fits into the overall RAG pipeline.
2. Bi-Encoders vs. Cross-Encoders: A Tale of Two Architectures
To understand why cross-encoders are so effective for reranking, we need to compare their architecture to the bi-encoders we've already used.
The Bi-Encoder Architecture (Independent Processing)
As we saw in the previous lesson, a bi-encoder processes the query and the documents independently.
- It creates a vector embedding for the query.
- It creates vector embeddings for all documents (this can be done offline).
- It then calculates a similarity score (like cosine similarity) between the query embedding and each document embedding.
This is computationally efficient because the document embeddings are pre-computed. At query time, we only need to embed the query and perform a fast vector search. However, this independence is also its weakness. The model never sees the query and the document at the same time, so it can miss subtle relationships.
The Cross-Encoder Architecture (Joint Processing)
A cross-encoder, on the other hand, processes the query and a document simultaneously.
- The query and a document are concatenated into a single input sequence, separated by a special
[SEP]token:[CLS] query_text [SEP] document_text [SEP]. - This combined input is fed through a Transformer model (like BERT).
- Because they are processed together, the self-attention mechanism can create connections between the tokens of the query and the tokens of the document.
- The model outputs a single score (a logit) that represents the relevance of the document to the query. This is not a similarity score but a direct prediction of relevance.

This joint processing allows the cross-encoder to capture much more nuance, such as negation, context, and word order, leading to a far more accurate relevance assessment.
The following reading provides an excellent, concise explanation of these differences and the trade-offs involved.
A Deep Dive into Cross-Encoders and How They Work
The article 'A Deep Dive into Cross-Encoders and How They Work' by Ranjan Kumar clearly explains the fundamentals of cross-encoders and their advantages.
Read sections 2, 3, and 9 of the article. They are titled 'What Are Cross-Encoders?', 'Why Cross-Encoders Matter...', and 'Strengths vs. Limitations'. This will solidify your understanding of the architecture, its pros and cons, and where it fits in a RAG system.
Test your understanding!
You have a corpus of 100,000 documents. A user submits a query. Why would it be computationally infeasible to use a cross-encoder for the initial retrieval stage (i.e., finding the top 25 candidates from the entire corpus)?
Show answer
A cross-encoder requires a full forward pass of a Transformer model for each query-document pair. To search the entire corpus, you would need to perform 100,000 forward passes, which would be extremely slow (taking minutes or even hours). In contrast, a bi-encoder pre-computes the 100,000 document embeddings. At query time, it only needs one forward pass for the query and then performs a very fast vector similarity search over the pre-computed embeddings. This is why cross-encoders are reserved for reranking a small number of candidates retrieved by a faster method.
3. Implementing the Reranking Step
Now for the practical part. We will add a Reranker component to our RAG pipeline from the previous lesson. Given your software engineering background, you'll appreciate how our modular design makes this integration straightforward.
We will use the sentence-transformers library, which provides a simple CrossEncoder class.
The Reranking Workflow
The new workflow inside our pipeline will look like this:

- Retrieve More Documents: We'll configure our
DenseRetrieverto fetch a larger number of documents, sayk=25, to ensure we have a good pool of candidates for the reranker (prioritizing recall). - Create Pairs: For the
top_kdocuments retrieved, we'll create(query, document_text)pairs. - Predict Scores: We'll pass this list of pairs to the
cross_encoder.predict()method. - Sort and Filter: We'll sort the documents based on the new cross-encoder scores and take the new top
n(e.g., top 5) to pass to the generator (prioritizing precision).
Let's see this in action. The following video and blog post provide excellent, practical examples of implementing this retrieve-and-rerank logic.
Do Reranking Models Actually Improve RAG?
Let's return to the 'Do Reranking Models Actually Improve RAG?' video. This segment provides a detailed walkthrough of the cross-encoder's mechanism and a code demonstration.
Watch from 09:40 to 14:46. This part explains the cross-encoder's inner workings with attention and then shows a clear Python implementation using sentence-transformers. Observe how the model's output is a raw score and how it effectively ranks passages about meditation.
The next resource provides a complete, end-to-end code example that you can study.
Sentence Embeddings. Cross-encoders and Re-ranking
The blog post 'Sentence Embeddings. Cross-encoders and Re-ranking' by Omar Sanseviero contains a fantastic, detailed implementation of the full retrieve-and-rerank pipeline.
Read the entire 'Retrieve and re-rank' section. It walks through loading a dataset, using a bi-encoder for initial retrieval, and then using a cross-encoder to rerank the results. Pay close attention to the code blocks showing how to call the bi_encoder and cross_encoder and how the final ranking changes.
4. Updating Our RAG Pipeline
Now, let's integrate this reranking logic into the Python code from our previous lesson. We'll add a Reranker class and modify the RAGPipeline to include this new step.
We need to install one library:
pip install sentence-transformers
Here is the updated code. Notice the new Reranker class and the changes in the RAGPipeline.query method.
import numpy as np
import faiss
from sentence_transformers import SentenceTransformer, CrossEncoder
import ollama
# --- Component 1: The Retriever (Unchanged from last lesson) ---
class DenseRetriever:
def __init__(self, documents, model_name='all-MiniLM-L6-v2'):
self.documents = documents
self.model = SentenceTransformer(model_name)
self.index = None
self._build_index()
def _build_index(self):
# We use numpy arrays for FAISS
embeddings = self.model.encode(self.documents, convert_to_numpy=True)
faiss.normalize_L2(embeddings) # Normalize vectors for cosine similarity with IndexFlatIP
embedding_dim = embeddings.shape[1]
self.index = faiss.IndexFlatIP(embedding_dim) # Using Inner Product for cosine similarity
self.index.add(embeddings)
def search(self, query, k=3):
query_embedding = self.model.encode([query])
faiss.normalize_L2(query_embedding)
distances, indices = self.index.search(query_embedding, k)
return [self.documents[i] for i in indices[0]]
# --- Component 2: The Reranker (New) ---
class Reranker:
def __init__(self, model_name='cross-encoder/ms-marco-MiniLM-L-6-v2'):
self.model = CrossEncoder(model_name)
def rerank(self, query, documents):
pairs = [(query, doc) for doc in documents]
scores = self.model.predict(pairs)
# Combine documents with their scores and sort
doc_score_pairs = list(zip(documents, scores))
doc_score_pairs.sort(key=lambda x: x[1], reverse=True)
# Return just the reranked documents
reranked_docs = [doc for doc, score in doc_score_pairs]
return reranked_docs
# --- Component 3: The Generator (Unchanged from last lesson) ---
class Generator:
def __init__(self, model='llama3:8b-instruct'):
self.model = model
self.prompt_template = """
You are a helpful assistant. Use ONLY the following pieces of retrieved context to answer the question.
If you don't know the answer, just say that you don't know. Don't try to make up an answer.
Context:
{context}
Question:
{question}
Answer:
"""
def generate_answer(self, query, retrieved_docs):
context_str = "\n---\n".join(retrieved_docs)
prompt = self.prompt_template.format(context=context_str, question=query)
response = ollama.chat(
model=self.model,
messages=[{'role': 'user', 'content': prompt}]
)
return response['message']['content']
# --- Component 4: The Enhanced RAG Pipeline ---
class RAGPipeline:
def __init__(self, retriever, reranker, generator):
self.retriever = retriever
self.reranker = reranker
self.generator = generator
def query(self, query_text, retrieve_k=10, rerank_k=3):
print(f"--- Querying for: '{query_text}' ---")
# 1. Retrieve a larger set of documents (High Recall)
retrieved_docs = self.retriever.search(query_text, k=retrieve_k)
print(f"\n--- Retrieved Top {retrieve_k} Documents (before reranking) ---")
for i, doc in enumerate(retrieved_docs):
print(f"Doc {i+1}: {doc[:100]}...") # Print snippet
# 2. Rerank the retrieved documents (High Precision)
reranked_docs = self.reranker.rerank(query_text, retrieved_docs)
# Keep only the top 'rerank_k' for the generator
final_docs = reranked_docs[:rerank_k]
print(f"\n--- Top {rerank_k} Reranked Documents (sent to LLM) ---")
for i, doc in enumerate(final_docs):
print(f"Doc {i+1}: {doc[:100]}...")
# 3. Generate an answer based on the reranked documents
answer = self.generator.generate_answer(query_text, final_docs)
return answer
# --- Example Usage ---
if __name__ == "__main__":
# A slightly more complex corpus to showcase the power of reranking
corpus = [
"The capital of France is Paris, a city famous for art and culture.",
"The Eiffel Tower, located in Paris, is a global cultural icon of France.",
"French cuisine is renowned for its wines, cheeses, and pastries.",
"In contrast to France, Berlin is the capital of Germany.",
"Paris Saint-Germain is a professional football club based in Paris.",
"The Louvre Museum in Paris is the world's largest art museum.",
"While beautiful, Paris has notable issues with air pollution.",
"The 2024 Summer Olympics will be held in Paris, France.",
"What is the population of Paris? The city proper has over 2 million residents.",
"Unlike Paris, London is the capital of the United Kingdom."
]
# 1. Initialize components
retriever = DenseRetriever(documents=corpus)
reranker = Reranker()
generator = Generator()
# 2. Initialize the pipeline
rag_pipeline = RAGPipeline(retriever=retriever, reranker=reranker, generator=generator)
# 3. Ask a question where reranking can make a difference
question = "What is the capital of France and what are its problems?"
final_answer = rag_pipeline.query(question, retrieve_k=5, rerank_k=2)
print("\n--- Final Answer ---")
print(final_answer)
When you run this code, pay close attention to the output. The initial retrieval might pull in documents that are broadly about "Paris" or "France". The reranker, however, will prioritize documents that specifically mention both the capital and its "problems," providing a much better context for the generator to formulate a precise answer.
Conclusion
Congratulations! You have now implemented one of the most impactful upgrades for a RAG system. By adding a cross-encoder reranker, you've created a more robust two-stage retrieval process that balances speed and accuracy.
Key Takeaways:
- Cross-encoders offer superior precision by jointly processing a query and a document, allowing for deep, token-level interaction via self-attention.
- They are computationally expensive, making them unsuitable for first-pass retrieval but perfect for a reranking stage on a small set of candidate documents.
- A retrieve-and-rerank architecture is a powerful pattern in modern search and RAG systems, combining a high-recall retriever (like a bi-encoder) with a high-precision reranker (a cross-encoder).
- Implementing this in a modular pipeline is straightforward and significantly improves the quality of the context provided to the generator, leading to more accurate and relevant final answers.
Preview of the Next Lesson:
We have just significantly improved the second stage of our retrieval process. But what about the first stage? Can we do better than using a generic, pre-trained bi-encoder? The answer is yes. In the next lesson, we will explore how to apply contrastive learning to train high-quality sentence embedding models (bi-encoders). This will allow us to create a custom retriever that is fine-tuned for the specific domain of our documents, further boosting the performance of our entire RAG pipeline.