Hello! Welcome to your sixth and final lesson in the "Agentic AI Systems" module.
In our last lesson, we designed agents that could create sophisticated plans by decomposing complex tasks into smaller, manageable steps. You learned about patterns like Plan-and-Execute, REWOO, and the LLM Compiler, which allow an agent to strategize its approach to a problem.
However, a plan is merely a hypothesis about the best way to achieve a goal. What happens when an action doesn't produce the expected result, or when the initial plan has flaws? Today, we address this by equipping our agents with the ability to learn from their experience. Your learning outcome is to implement agent reflection and self-correction mechanisms for improvement.
This is the bridge from a static planner to a dynamic learner—an agent that doesn't just execute, but also introspects, critiques, and improves.

1. The Core Idea: The Reflective Loop
Self-reflection allows an agent to analyze its own actions and outputs, identify flaws or areas for improvement, and use that critique to generate a better result in the next iteration. In essence, we are asking the LLM to act as its own code reviewer or editor.
This process introduces a feedback loop into the agent's architecture, moving beyond a simple Query -> Action sequence.
To start, let's formally define self-reflection in the context of autonomous agents. Lilian Weng's foundational blog post provides a concise introduction.
Read the short introductory paragraph in the 'Self-Reflection' section. This will frame the concept as a mechanism for iterative improvement by refining past decisions.
The simplest form of reflection can be visualized as a cycle: Generate -> Reflect -> Refine. An agent produces an initial output, another part of the agent (or the same LLM with a different prompt) critiques that output, and this critique is then fed back into the generator to produce an improved version.
Reflection Agent From Scratch | Agentic Patterns Series
Let's watch a clear explanation of this fundamental 'generate-reflect' loop.
Watch the segment from 01:14 to 04:43. The speaker clearly breaks down the two main blocks, 'Generate' and 'Reflect', and explains the iterative workflow. Pay attention to the two ways the loop can be terminated: a fixed number of steps or a 'stop' signal from the reflection block.
This simple loop is the foundation for all the more complex self-correction architectures we will explore.

2. A Spectrum of Reflection Architectures
While the core loop is simple, its implementation can range from a basic two-step process to a sophisticated, multi-agent system that can search for new information. Let's explore this spectrum.
2.1. Basic Reflection
This is the direct implementation of the generate-reflect loop. An LLM generates a response, and then a second LLM call (or the same LLM with a new prompt) is made to critique it. This critique is then used to generate a revised response.
Breaking Down & Testing FIVE LLM Agent Architectures - (Reflexion, LATs, P&E, ReWOO, LLMCompiler)
Let's review the basic reflection pattern and, more importantly, understand its primary limitation.
Watch the segment from 01:10 to 06:33. The video explains the basic reflection loop used to write an essay. Critically, notice the speaker's conclusion: without access to external tools, the agent's 'reflections' are not grounded in reality and can lead to confident hallucinations (e.g., made-up statistics).
This highlights a crucial principle: reflection is most powerful when it can identify and fix knowledge gaps. A simple reflective loop can improve style and structure, but to improve factual accuracy, the agent needs tools to find external information.
2.2. The Reflexion Framework: Reflection with Tool Use
The Reflexion framework (from the paper by Shinn et al.) formalizes this idea of tool-augmented reflection. Here, the reflection step doesn't just produce a critique; it produces an actionable plan to address the critique, often involving tool use.
Building a Self-Correcting AI: A Deep Dive into the Reflexion Agent
To understand the theory behind this powerful framework, we'll read a few sections from a deep-dive article that explains its core philosophy.
Read the following three sections: Introduction: The Problem with “System 1” AI: This frames reflection as a move towards more deliberate, 'System 2' thinking. The Core Philosophy: Reinforcement through Language, Not Weights: Understand the concept of 'verbal reinforcement' and the 'semantic gradient signal'. This is key—we are guiding the LLM with text, not retraining it. Episodic Memory: The Key to Iterative Learning: Grasp the importance of storing past reflections to learn from mistakes over time.
The workflow for a Reflexion agent is more sophisticated than the basic loop:
- Generate: Produce an initial response.
- Reflect & Plan: Critique the response and generate a list of search queries or tool calls to find missing information or verify facts.
- Act: Execute the tool calls.
- Revise: Synthesize the original response, the critique, and the new information from the tools into a final, improved output.
Breaking Down & Testing FIVE LLM Agent Architectures - (Reflexion, LATs, P&E, ReWOO, LLMCompiler)
Let's see a breakdown of this more advanced workflow.
Watch from 06:33 to 12:42. This segment explains the Reflexion architecture. Contrast this flow with the basic reflection loop. Notice how the 'tool execute' step is now a distinct and crucial part of the process, allowing the agent to ground its revisions in external data.
2.3. Advanced Reflection: Language Agent Tree Search (LATS)
Taking this a step further, what if an agent could explore multiple paths of reflection and revision at once? This is the idea behind LATS, which uses principles from Monte Carlo Tree Search.
Instead of generating one reflection, the agent generates multiple candidate thoughts or actions. It then uses another LLM call to "score" these candidates, pursues the most promising path, and repeats the process, effectively searching a tree of possible reasoning paths. Your CS background will recognize this as a search algorithm applied to the semantic space of ideas.
Breaking Down & Testing FIVE LLM Agent Architectures - (Reflexion, LATs, P&E, ReWOO, LLMCompiler)
For a glimpse at the state-of-the-art, let's briefly look at the LATS architecture.
Watch from 12:42 to 21:01. You don't need to memorize the details, but understand the core concept: instead of a linear revision process, LATS generates and scores multiple candidates at each step, turning reflection into a search problem. Note the potential pitfall: the LLM used for scoring can be unreliable and tend to overscore.
3. Implementation Blueprint
Now, let's translate the Reflexion theory into a practical implementation blueprint. We'll use the architecture laid out in the "Deep Dive" article, which is a perfect fit for your software engineering background due to its modular design.
3.1. Architecture: Responder, Search Tool, and Revisor
The agent is broken down into three distinct components:
- The Responder: The initial actor that generates a first-draft answer and, crucially, a self-critique with a plan (search queries) to improve it.
- The Search Tool: An external tool (e.g., a web search API) that executes the plan from the Responder.
- The Revisor: A second actor that takes the original query, the first draft, the critique, and the new search results to synthesize a final, cited answer.
3.2. Implementation Trick: Pydantic for Structured Reflection
A key challenge is reliably getting the critique and search queries from the LLM. A brilliant and robust technique is to use Pydantic models to define the desired JSON output structure. The descriptions in the Pydantic fields act as targeted sub-prompts, guiding the LLM to produce the exact output we need.
Building a Self-Correcting AI: A Deep Dive into the Reflexion Agent
This technique of using Pydantic for structured output is a cornerstone of reliable agent development. Let's see how it's used to fuse the generator and reflector roles into a single, efficient LLM call.
Read the sections 'The Foundation: Pydantic for Structured Outputs', 'Step 1: The First Responder', and 'Step 3: The Revision Cycle'. Focus on: How the description fields in the Reflection Pydantic model guide the LLM's critique. How the first_responder chain is constructed, forcing the LLM to use the AnswerQuestion tool (our Pydantic model). How the revisor chain uses a different instruction set and the ReviseAnswer model to incorporate new information and generate citations.
Test your understanding!
Why is using a Pydantic model with llm.bind_tools more robust for implementing reflection than simply asking the LLM to "critique your answer and provide search queries in a JSON format" in the prompt?
Show answer
While simple prompting might work, it's brittle. The LLM could fail to generate valid JSON, omit a field, or add extra conversational text. By using bind_tools with a Pydantic model, we leverage the model's function-calling capabilities, which are specifically fine-tuned to produce structured, valid JSON that conforms to a provided schema. This makes parsing the output virtually foolproof and the overall system much more reliable.
3.3. Orchestration with LangGraph
The cyclical nature of reflection—respond, search, revise, and potentially loop back—is a perfect use case for LangGraph, which you saw in the previous lesson.
Building a Self-Correcting AI: A Deep Dive into the Reflexion Agent
Finally, let's see how to wire these components together into a stateful, looping agent.
Read 'Part 4: Orchestration — Bringing It All Together with LangGraph'. Focus on the structure: AgentState: The TypedDict that defines the agent's working memory. Nodes: Functions like call_responder, call_search_tool, and call_revisor that wrap our chains. Conditional Edge: The decide_to_finish function that controls the loop, deciding whether to end or continue with another revision cycle.
3.4. Putting It All Together: A Code Example
To see this in a self-contained example, let's return to the first video, which implements a similar agent in a simple Python class.
Reflection Agent From Scratch | Agentic Patterns Series
This segment shows a complete, runnable implementation of a reflection agent, demonstrating the iterative refinement of a piece of Python code.
Watch from 06:43 to 18:09, skipping around as needed. You don't need to follow every line of code, but observe the process: The initial, simple generation of the merge sort algorithm. The 'Andrej Karpathy' persona generating a detailed critique. The structure of the ReflectionAgent class with generate, reflect, and run methods. The final execution showing the agent looping through steps, with the code getting progressively better (e.g., adding docstrings, complexity analysis, a main function).
Conclusion
You have now learned how to build agents that don't just act, but also think about their actions. This capacity for self-correction is a fundamental step toward creating more reliable, accurate, and intelligent AI systems.
Key Takeaways:
- Reflection is a Core Agentic Capability: It moves an agent from being a simple executor to a dynamic learner, enabling "System 2" thinking.
- A Spectrum of Architectures: Self-correction can be implemented as a simple Generate-Reflect-Refine loop, a tool-augmented Reflexion cycle, or an advanced LATS search.
- Grounding is Crucial: Effective reflection, especially for factual tasks, requires grounding the agent with external tools to verify information and fill knowledge gaps.
- Implementation Best Practices: Robust implementation relies on forcing structured outputs (e.g., with Pydantic), creating modular components (Responder, Revisor), and using a framework like LangGraph for stateful orchestration.
- Verbal Reinforcement: This entire process works through "verbal reinforcement"—using natural language critiques as a "semantic gradient" to guide the LLM's output, without needing costly model retraining.
Preview of the Next Module:
This lesson concludes our deep dive into agentic systems. You have now explored the core pillars: perception, planning, memory, and reflection. In our next module, "Multimodal and Cross-Domain AI," we will expand the agent's world. You will learn how to build systems that can not only process text but also understand images with Vision Transformers (ViT), process audio with models like Whisper, and generate speech, creating truly multimodal agents that can interact with the world in richer ways.