Hello! Welcome to the third lesson in our module on Agentic AI Systems.
In the last two lessons, we established the core architecture of an agent. You learned about the perception-action loop and, critically, implemented function calling to give your agent "hands"—the ability to act on the world by interacting with external tools and APIs.
However, the agent we've considered so far is reactive. It sees a query, decides to use a tool, uses it, and then summarizes the result. This is powerful but limited. It doesn't allow for multi-step problem-solving, planning, or self-correction. To build truly intelligent agents, we need to give them the ability to "think."
This lesson introduces a foundational technique to achieve that. Your learning outcome is to implement the ReAct (Reason + Act) prompting framework for agentic behavior. This framework synergizes reasoning and acting, enabling the agent to create plans, execute them, observe the outcomes, and adjust its strategy—much like a human would.
1. From Acting to Reasoning and Acting
The limitation of simple function calling is that the LLM's reasoning is implicit. It decides on a tool and returns a call, but it doesn't articulate why it chose that tool or what its broader plan is.
ReAct, a paradigm introduced in the paper "ReAct: Synergizing Reasoning and Acting in Language Models", explicitly combines the strengths of:
- Reasoning (like in Chain-of-Thought prompting)
- Acting (like in the function calling you implemented)
The core idea is to prompt the LLM to follow an interleaved sequence of Thought → Action → Observation.

Let's break down this powerful loop.
Inside the ReAct Agent: A Step-by-Step Breakdown
This blog post provides an excellent, detailed breakdown of the ReAct agent's core cycle. It will give you a solid conceptual foundation before we dive into the code.
Please read the entire section '2. Inside the ReAct Agent: A Step-by-Step Breakdown'. Pay close attention to the distinct roles of Thought, Action, PAUSE, and Observation. The full example at the end of the section clearly illustrates how these steps chain together to solve a complex query.
As you've just read, the loop works as follows:
- Thought đź§ : The LLM first thinks about the problem. It breaks the query down, forms a plan, reflects on previous steps, or decides what information it needs next. This is its "inner monologue."
- Action ⚡: Based on its thought, the LLM decides to take an action. This action is a function call to one of its available tools.
- Observation đź‘€: The application code executes the action and returns the result to the agent. This result is the observation. The agent now knows the outcome of its action.
- Repeat: The agent takes this new observation into account and starts the loop again with a new thought, refining its plan until it has enough information to provide a final answer.
This approach is fundamentally different from a single function call. It creates a conversational, multi-turn process where the agent can tackle problems that require several steps, such as "What is the weather in the capital of the country that won the last FIFA world cup?"
2. Implementing ReAct from Scratch
Frameworks like LangChain have pre-built ReAct agents, but building one from scratch provides a much deeper understanding of the mechanics. Given your software engineering background, you'll be able to appreciate the interplay between the prompt, the model, and the control loop you'll write.
We will build an agent that can solve multi-step arithmetic problems involving fetching data. This will require us to implement the three key components of the ReAct system.
The Core Components:
- A Powerful System Prompt: This is the most critical part. The prompt must explicitly instruct the LLM to follow the Thought-Action-Observation format and provide examples.
- A Set of Tools: Simple Python functions the agent can call (e.g., a calculator, a data lookup function).
- An Orchestration Loop: A Python loop that manages the entire process: sending prompts, parsing the LLM's output for actions, executing them, and feeding the observations back.
The following video provides a complete walkthrough of building this entire system in Python, without relying on any agent-specific frameworks.
Python: Create a ReAct Agent from Scratch
This video by Alejandro AO is a complete guide to building a ReAct agent from the ground up. We will use it to guide our implementation step-by-step.
First, watch the conceptual overview from 01:27 to 10:18. The diagram he presents is an excellent visualization of the process. Then, continue from 10:18 to 15:01 to see how to set up the environment and test the basic LLM client. You can use any model provider you like (e.g., OpenAI, Anthropic) as long as it's a capable instruction-following model.
Step 1: The Agent Class and System Prompt
First, we'll define a simple Agent class to encapsulate the state (like the message history) and the logic for calling the LLM. Then, we will craft the all-important system prompt.
The system prompt is the "operating system" for the agent. It defines the rules of engagement, the format for thoughts and actions, and lists the available tools with descriptions.
Python: Create a ReAct Agent from Scratch
Let's continue with the video to implement the Agent class and the system prompt.
Watch from 15:01 to 29:12. First, see how the Agent class is structured. Notice how it maintains a list of messages, which serves as its short-term memory. Then, pay very close attention to how the system_prompt is constructed. It includes the rules, the format, a description of the tools (calculate, get_planet_mass), and a few-shot example of a complete Thought-Action-Observation trajectory.
Test your understanding!
In the ReAct system prompt, why is it crucial to include a complete, step-by-step example of a question being solved (e.g., the What is the mass of the Earth times 2? example in the video)? What role does this example play?
Show answer
This is a form of few-shot prompting. The example serves as a concrete template for the LLM to follow. It demonstrates the exact output format expected (Thought: ..., Action: ..., PAUSE, Final Answer: ...). By seeing a valid trajectory, the model learns to replicate that pattern for new, unseen questions. Without it, the model might not reliably adhere to the strict, interleaved reasoning and acting structure required by the ReAct framework.
Step 2: The Orchestration Loop
We now have an agent that can produce one step of thought and action. The final piece is to write a control loop in our application that automates the cycle:
- Get the agent's response.
- If it contains an
Action, parse it. - Execute the action's corresponding tool.
- Create an
Observationstring with the result. - Feed this observation back to the agent and repeat.
- If the response contains
Final Answer, break the loop.
This loop is the orchestrator that brings the agent to life, allowing it to autonomously perform multiple steps to solve a problem. Your experience with control flow and string parsing in Python will be very useful here.
Python: Create a ReAct Agent from Scratch
This final video segment walks through the implementation of this orchestration loop, bringing everything together into a fully autonomous agent.
Watch from 37:43 to 53:43. This is the culmination of our work. Follow how a while loop is used to manage the iterations. Pay attention to the logic for parsing the Action string (using regular expressions) to extract the tool name and arguments, executing the tool, and continuing the loop with the new observation.
After completing this, you will have a fully functional ReAct agent that you built from scratch. You can test it with complex queries like "What is the mass of Jupiter plus the mass of Earth, all divided by 2?" and watch it reason and act its way to the solution.
Conclusion
Today, you've taken a massive step forward in building agentic systems. You've moved beyond simple tool use and implemented the ReAct framework, giving your agent an "inner monologue." This allows it to tackle complex, multi-step tasks that require planning, information gathering, and dynamic adjustments.
Key Takeaways:
- ReAct = Reason + Act: It combines chain-of-thought reasoning with tool use in an interleaved
Thought -> Action -> Observationloop. - The System Prompt is Critical: It defines the agent's entire behavior, forcing it to follow the ReAct structure, detailing available tools, and providing few-shot examples.
- The Application is the Orchestrator: The LLM only generates the thoughts and action plans. Your Python code is responsible for the control loop that parses actions, executes tools, and feeds observations back to the model.
- Building from scratch reveals the underlying mechanics that are often abstracted away by frameworks, giving you a robust mental model of how these agents operate.
Preview of the Next Lesson:
The agent we built today has a simple form of memory—its conversation history. However, as conversations get longer, this history can exceed the LLM's context window. In our next lesson, we will design and implement basic memory systems for agents, exploring strategies to manage and summarize conversation history to enable long-running, stateful interactions.