Hello! Welcome to your fourth lesson on Agentic AI Systems.
In our last session, you masterfully implemented the ReAct framework, giving your agent an "inner monologue" to reason through multi-step tasks. The agent we built used a simple list of messages as its memory. This is effective for a single conversation but has two major limitations:
- It's ephemeral: The memory is lost once the session ends.
- It's finite: As the conversation grows, it will eventually exceed the LLM's context window.
Today, we will address these limitations head-on. Your learning outcome is to design and implement basic memory systems for agents. We will explore how to give our agents both a short-term working memory and a persistent long-term memory, enabling them to learn and retain information across multiple interactions.
1. A Conceptual Framework for Agent Memory
Before we write any code, it's crucial to have a solid mental model for what "memory" means for an agent. Drawing inspiration from cognitive psychology can be incredibly helpful. An agent's memory can be broadly categorized into several types, each serving a different purpose.
Building Brain-Like Memory for AI | LLM Agent Memory Systems
This video from Adam Lucek provides an excellent overview of different memory types for AI agents, drawing parallels to human cognition. It will help us establish a strong conceptual foundation.
Please watch the following segments to understand the different categories of memory: Introduction (00:00 - 03:15): Understand why memory is crucial and get an overview of the four memory types. Working Memory (03:15 - 05:45): This is the agent's 'short-term memory' or active consciousness. Note how it maps directly to the conversation history we've already used. Episodic Memory (08:32 - 10:45): This is the memory of past experiences or 'episodes'. Focus on the idea of storing not just the conversation but also reflections or takeaways. Semantic Memory (21:56 - 24:10): This is the agent's knowledge base of facts about the world. Understand its connection to Retrieval-Augmented Generation (RAG). Procedural Memory (30:11 - 33:00): This is the 'how-to' memory. Note that for our purposes, this is often embedded in the agent's code or model weights, so we'll focus on the other three.
As you've just seen, we can think of an agent's memory system as having several layers:

- Working Memory (Short-Term): The context of the current interaction. This is typically the history of messages in the current conversation. It's volatile and fast.
- Long-Term Memory: A persistent store of information that survives across sessions. We can further divide this into:
- Episodic Memory: Recalls past conversations and interactions. Answers "What happened between us before?"
- Semantic Memory: Stores general knowledge and facts. Answers "What is true about the world or the user?"
For this lesson, we will focus on implementing a robust Working Memory and a foundational Semantic Memory that can be updated, laying the groundwork for more complex episodic systems later.
2. From Simple History to a True Memory System
The simplest form of memory, which you've already implemented, is just keeping a running list of user and assistant messages.
Building an AI agent from scratch in Python
This article provides a clear, simple implementation of a basic short-term memory system. It's a great recap of the starting point we're building upon.
Read the section 'Component 2: (Conversation) Memory'. Notice how adding a messages list to the Agent class immediately gives it the ability to reference prior turns in the conversation. This is the core of working memory.
The challenge, as mentioned, is the limited context window. A simple but effective technique to manage this is Context Compaction. Before the history becomes too large, you can use an LLM to summarize the older parts of the conversation, replacing many messages with a single summary message.
# Conceptual example of context compaction
def compact_context(conversation_history: list, max_tokens: int) -> list:
# (Implementation would use a tokenizer to count tokens)
if num_tokens(conversation_history) < max_tokens:
return conversation_history
# Keep the most recent turns intact
recent_messages = conversation_history[-6:]
old_messages = conversation_history[:-6]
# Use an LLM to summarize the old messages
summary_prompt = f"Summarize the following conversation concisely: {old_messages}"
summary = call_llm(summary_prompt)
# Replace the old messages with the summary
return [{"role": "system", "content": f"Summary of prior conversation: {summary}"}] + recent_messages
This is a powerful pattern for managing working memory in long-running conversations.
3. Implementing a Persistent, Updatable Long-Term Memory
Now for the main event: creating a memory that persists. We'll use a vector database to build a semantic memory store, but with a crucial addition—the ability to update and delete memories to prevent contradictions.
This process involves four key challenges:
- Extraction: Deciding what information is worth remembering.
- Indexing: Storing that information for efficient retrieval.
- Retrieval: Fetching relevant memories when needed.
- Updating: Modifying or deleting stored memories when new information supersedes them.
We will follow a practical, from-scratch implementation inspired by the Mem0 paper. The video below uses the DSPy framework, but the underlying logic is pure Python and can be implemented with standard libraries and an LLM client. Your experience with Python and APIs will allow you to easily adapt these concepts.
Step 1: Extraction and Indexing
First, we need to extract atomic "memories" from a conversation and store them in a vector database for semantic search.
How to build your own long-term Agentic Memory System for LLMs | Mem0 from scratch in DSPy
This video by 'Neural Breakdown with AVB' demonstrates building a complete agentic memory system. We'll start with how to extract and index memories.
Watch from 12:12 to 27:00. Focus on these core concepts: Memory Extraction (12:12 - 20:29): Observe how an LLM is prompted to extract specific, structured pieces of information (factoids) from a conversation transcript. The use of Pydantic models to define the output structure is a great pattern you can replicate. Indexing (20:29 - 27:00): Understand the workflow: take the extracted text, generate an embedding using a model like OpenAI's text-embedding-3-small, and then insert the embedding along with its payload (the text, user ID, categories, etc.) into a vector database like Qdrant.
At this stage, you have a RAG pipeline: text is converted into memories, embedded, and stored. When a new query arrives, you embed the query, search the vector DB for similar memories, and add them to the agent's context.
Step 2: The Critical Update Mechanism
This is what transforms a simple RAG store into a dynamic memory system. What happens when information changes?

To implement this, we'll give our agent a new set of tools specifically for memory management. This perfectly builds on the ReAct and function-calling skills you developed in the last two lessons.
How to build your own long-term Agentic Memory System for LLMs | Mem0 from scratch in DSPy
Let's continue with the video to implement the most complex and powerful part of our memory system: the update loop.
Watch from 32:29 to 40:41. This section is dense but crucial. Pay close attention to how a ReAct agent is used to manage the memory itself. The Tools: Notice the Python functions add_memory, update_memory, and delete_memory. These are the 'actions' the memory-management agent can take. The Logic: The agent is given a new piece of information and similar memories retrieved from the database. Its task is to 'reason' and 'act' by calling one of the tools to keep the memory store consistent and accurate. Implementation: The agent calls the update tool with the ID of the old memory and the new text. Your Python code then handles the database operation (e.g., deleting the old vector and inserting a new one).
Test your understanding!
Why is it generally a bad idea to directly update the payload (text content) of an existing point in a vector database without also updating its vector embedding?
Show answer
The vector embedding is a numerical representation of the semantic meaning of the text. If you change the text (e.g., from "user likes coffee" to "user hates coffee"), the meaning changes completely. However, the vector stored in the database still represents the old meaning. This creates a mismatch. A future search for "user's beverage preferences" might still retrieve this memory based on the old vector, but the agent will read the new, conflicting text. The correct approach, as shown in the video, is to delete the old point entirely and insert a new point with the updated text and its corresponding new embedding.
Step 3: Integrating Memory into the Agent Loop
Now, we integrate this memory system into our main agent's operational cycle.
How to build your own long-term Agentic Memory System for LLMs | Mem0 from scratch in DSPy
This final segment shows how all the pieces come together to create a chatbot with a persistent, dynamic memory.
Watch from 40:41 to 51:21. This demonstrates the full loop in action. Follow how the agent: Responds to a query, but first fetches similar memories to enrich its context. When new information is introduced (e.g., "I like to sleep"), it triggers the memory update agent in the background to add this new fact. When conflicting information is given (e.g., "I hate to sleep"), it correctly calls the update function to modify the existing memory rather than adding a new, contradictory one.
By completing this, you've designed and implemented a full-fledged memory system that can extract, index, retrieve, and intelligently update information over time.
Conclusion
In this lesson, you've moved beyond the simple, ephemeral memory of a conversation list and designed a truly dynamic memory system for your agent. This is a fundamental component for building sophisticated, long-running AI assistants that learn and adapt.
Key Takeaways:
- Agent memory is layered: We can model it with a Working Memory (current context) and a Long-Term Memory (semantic and episodic).
- Working Memory can be managed: Techniques like context compaction (summarization) help overcome context window limitations.
- Long-Term Memory needs to be dynamic: A simple vector store is not enough. A robust system must handle updates, preventing the accumulation of stale or contradictory information.
- LLMs can manage their own memory: By using a ReAct agent with tools like
add,update, anddelete, you can create a self-managing memory system that intelligently maintains its own knowledge base.
Preview of the Next Lesson:
With a robust memory system in place, your agent can now remember goals, sub-tasks, and the results of its actions over extended periods. This capability is the bedrock of advanced planning. In the next lesson, we will apply task decomposition and planning for multi-step agent workflows, enabling your agent to tackle highly complex problems by breaking them down into smaller, manageable steps and tracking its progress using its newly acquired memory.