Skip to main content
Create your own
Lesson illustration

Building a Basic RAG Workflow

Hello! Welcome back to our exploration of AI and LLM integrations in n8n.

In our last lesson, we unpacked the theory behind Retrieval-Augmented Generation (RAG). You learned what RAG is, why it's essential for creating knowledgeable and fact-based AI, and the critical roles played by vector databases and embeddings.

Today, we move from theory to practice. You will apply those concepts to build your first basic RAG workflow in n8n. This involves two key parts: first, a workflow to populate a vector store with your custom knowledge, and second, an agent workflow to retrieve context from that store and answer questions. This is a foundational skill for creating powerful, custom AI solutions.

Let's get building!

The Two Halves of a RAG System

A complete RAG system in n8n is best understood as two separate but connected workflows:

  1. Ingestion Workflow: This workflow's job is to read your source documents (like text files, PDFs, etc.), break them into manageable chunks, convert those chunks into vector embeddings, and store them in a vector database. You typically run this workflow whenever your knowledge base needs to be updated.
  2. Query Workflow: This is the "live" part of the system. It's usually an AI agent that, when it receives a user's question, retrieves the most relevant chunks of information from the vector store and uses them as context to generate an answer.
n8n RAG Workflow Example
This image shows a perfect visual representation of the two workflows we are about to build: one for loading data into a vector store, and another for chatting with it.

For this lesson, we will use the following stack, which is common for getting started:

  • Data Source: Google Drive, to hold our text files.
  • Vector Store: Pinecone, a popular and easy-to-use managed vector database.
  • Embedding & Chat Model: OpenAI, for both creating the embeddings and generating the final answer.

Part 1: The Ingestion Workflow - Populating the Knowledge Base

Our first task is to create a workflow that reads documents from Google Drive and populates our Pinecone vector database.

Step 1.1: Preparing Your Data and Database

Before we open n8n, we need our source data and a database to put it in.

N8N Tutorial: Creating a RAG Agent in n8n for Beginners! (Full Guide)

The AI Foundations channel provides an excellent end-to-end guide for building this RAG agent. We'll use it to construct our workflows. First, let's get our data and vector database ready.

Watch from 04:15 to 07:50. This will guide you through: Organizing your source data: Create a folder in your Google Drive and upload a few .txt files with some sample content. This will be your knowledge base. Setting up Pinecone: Create a free account on Pinecone and set up a new index. Pay close attention when the video specifies the embedding model (text-embedding-3-small) and its dimensions (1536). This information is crucial, as the model and its dimensions must be consistent throughout your workflow.

Step 1.2: Building the Ingestion Workflow in n8n

Now that you have your source files and an empty Pinecone index, let's build the n8n workflow to connect them. This process involves fetching file metadata, downloading the file content, and then using the Vector Store node to handle chunking, embedding, and uploading.

N8N Tutorial: Creating a RAG Agent in n8n for Beginners! (Full Guide)

Let's build the workflow to connect your Google Drive to Pinecone. This part of the video walks through the entire ingestion process, node by node.

Watch carefully from 07:50 to 25:30. This is the longest and most detailed part of the lesson. The video will show you how to: Fetch file metadata from Google Drive. Download the file content into n8n. Configure the Pinecone Vector Store node to insert the documents. This is the key step where you will: Connect your OpenAI credentials for the embedding model. Set up the document loader to process the downloaded files. Configure the text splitter to chunk your documents. Since your background is in software development, the process of setting up OAuth credentials for Google Drive (around 09:50-15:30) will be familiar in principle, though the specifics are important. Take your time, pause the video, and build along in your own n8n canvas.

After successfully running this workflow, you can check your Pinecone index. You should see that it's no longer empty and now contains a number of vectors far greater than the number of files you uploaded. This is because the text splitter broke your documents into many smaller chunks.

Part 2: The Query Workflow - Chatting With Your Data

With our knowledge base populated, it's time to build the agent that can use it. This second workflow will receive a user's question, retrieve relevant context from Pinecone, and use an LLM to generate an answer.

N8N Tutorial: Creating a RAG Agent in n8n for Beginners! (Full Guide)

Now for the fun part. Let's build the chat agent that can access your newly created knowledge base.

Watch from 25:30 to 31:03. This section covers building the agent itself. You will learn to: Start a workflow with the On chat message trigger. Add and configure the AI Agent node. Add the Vector Store Question Answer tool and connect it to your Pinecone index. Critically, you will configure the same embedding model (text-embedding-3-small) that you used for ingestion. This ensures the question vector and the document vectors can be compared correctly.

Once this workflow is set up, you can use the Chat button in the n8n UI to test it. Ask a specific question that can only be answered using the content from one of your text files.

N8N Tutorial: Creating a RAG Agent in n8n for Beginners! (Full Guide)

Let's see the result of our work. The video demonstrates how to interact with the agent and how it retrieves information to form an answer.

Watch the final segment from 31:03 to 34:26 to see the RAG agent in action.

You have now built a complete, end-to-end RAG system!

Test your understanding!

In the agent workflow, why is it absolutely critical that the embedding model used in the "Vector Store Question Answer tool" is identical to the one used in the ingestion workflow's "Pinecone Vector Store" node? What would happen if you used two different models (e.g., OpenAI for ingestion and Cohere for retrieval)?

Show answer

Different embedding models represent text in different high-dimensional "vector spaces." The entire principle of semantic search relies on comparing the vector of the query with the vectors of the documents to find the "closest" matches. This comparison is only meaningful if all vectors exist in the same coordinate system.

If you used two different models, the query vector and the document vectors would be in completely different, incompatible vector spaces. The similarity search would produce meaningless results, as there would be no logical concept of "closeness" between them. The agent would fail to retrieve any relevant context and would likely fall back to its base knowledge, unable to answer questions about your documents.

An Important Note on Production-Ready Workflows

Our ingestion workflow is great for adding new documents, but it has a limitation: it doesn't handle updates to existing documents. If you modify a file in Google Drive and re-run the ingestion workflow, you will create duplicate chunks in your vector store, leading to redundant and potentially conflicting information.

A more robust, production-ready system must handle updates gracefully. The standard pattern is to first delete any existing chunks associated with a file before inserting the new, updated chunks.

This RAG AI Agent with n8n + Supabase is the Real Deal

This video from Cole Medin demonstrates this exact update pattern using Supabase as the vector store. The principle is the same regardless of the database.

Watch the segment from 12:12 to 13:45. You don't need to build this right now, but focus on the concept. Notice how the workflow includes a step to 'delete all of the old vectors for this file' before inserting the new, updated content. This DELETE then INSERT logic is a key pattern for maintaining a clean and accurate knowledge base.

Conclusion

Congratulations! You've successfully built a functional RAG system in n8n. This is a massive step towards creating sophisticated AI applications that can reason about your specific data.

Key Takeaways:

  • A basic RAG system consists of two workflows: one for ingestion (reading, chunking, embedding, storing) and one for querying (embedding, retrieving, generating).
  • The Vector Store node in n8n is a powerful abstraction that combines document loading, text splitting, and embedding into a single interface.
  • The AI Agent uses a Tool (the Vector Store Question Answer tool) to connect to the knowledge base for retrieval.
  • Model consistency is crucial: The same embedding model must be used for both ingestion and retrieval.
  • Production-level RAG requires careful handling of data lifecycle, including updates and deletions, to avoid duplication.

Preview of the Next Lesson:

Now that our agent has both conversational memory (from a previous lesson) and a long-term knowledge base, we can make it do more complex things. In the next lesson, we will learn how to implement multi-step AI agent chains for complex reasoning tasks, enabling our agent to not just answer questions, but to execute sequences of actions to achieve a goal.

Can't find a good explanation? Sign up and we'll make it for you

Sign up