Hello! Welcome to the first lesson of our module on Agentic AI Systems.
In the previous module, we mastered Retrieval-Augmented Generation (RAG), creating systems where a Large Language Model (LLM) reasons over a set of provided documents to answer questions. This is a powerful pattern, but the LLM's world is limited to the context we give it.
Now, we'll take a significant leap forward. We will explore systems where the LLM doesn't just reason, but acts. These are AI agents: autonomous systems that perceive their environment, make decisions, and use tools to achieve goals. This lesson lays the foundation for this entire topic. Your objective is to learn how to design a basic agent using a perception-action loop.
1. What is an AI Agent?
At its core, an AI agent is a system that can perceive its environment and act upon that environment to achieve a specific goal. Think of it as an entity with a degree of autonomy. This is fundamentally different from a simple reactive model, like a standard LLM call, which just maps a single input to a single output without any broader sense of purpose or ability to perform multi-step actions.
The Agentic AI Handbook: A Beginner's Guide to ...
To understand this crucial distinction, let's start with 'The Agentic AI Handbook' from freeCodeCamp. This resource clearly contrasts agentic AI with reactive AI.
Please read the section titled 'Agentic vs Reactive AI'. Focus on how agentic systems can break down complex tasks and are proactive, unlike reactive systems.
Agents can be classified based on their complexity and capabilities. Let's watch a short video from IBM Technology that outlines a few common types.
5 Types of AI Agents: Autonomous Functions & Real-World Applications
This video, '5 Types of AI Agents', will introduce you to a spectrum of agent designs, from the very simple to the highly advanced.
Watch from the beginning to 02:52. This segment introduces the concept of an agent and details the simplest type: the simple reflex agent. Pay close attention to the diagram showing how it perceives the environment and follows 'condition-action' rules.
The simple reflex agent described in the video is the most basic form of an agent. It operates on a direct if-this-then-that logic. Our focus in this module will be on building more sophisticated agents that use LLMs for their decision-making, but they all share a fundamental operating principle: the perception-action loop.
2. The Perception-Action Loop
Every agent, from a simple thermostat to a complex self-driving car, operates on a continuous cycle known as the perception-action loop. This is the fundamental design pattern for any autonomous system.
The loop consists of two main phases:
- Perception: The agent uses sensors to gather information about the current state of its environment. This information is called a percept.
- Action: The agent processes the percepts and uses its internal logic to choose an action. It then executes this action using effectors (or actuators), which in turn changes the state of the environment.
This change is then observed in the next perception phase, and the cycle continues until the agent's goal is met.

For the agents we will build, the "environment" is often a digital one (like a computer's file system, a set of APIs, or a user interface), and the "brain" of the agent is an LLM.

3. Core Components of an LLM-Powered Agent
Translating the abstract perception-action loop into a modern, LLM-based system involves three key components. Your background in software engineering will make this structure feel very familiar, as it's analogous to a program that can call external services and maintain state.
- A Large Language Model (LLM): This is the agent's "brain" or reasoning engine. It's responsible for the "thinking" or "planning" part of the loop—understanding the goal, processing percepts, and deciding which action to take next.
- Tools: These are the agent's "hands" or effectors. A tool is essentially a function or an API that the agent can call to interact with its environment. Examples include a web search function, a calculator, a Python code interpreter, or an API for a vector database. The agent doesn't execute these directly; the LLM decides which tool to use and with what arguments, and an outer system executes it.
- Memory: This allows the agent to retain information from previous perception-action cycles. It can range from short-term memory (like the history of the current conversation) to long-term memory (a database of facts learned over time). Memory is crucial for handling multi-step tasks and learning from experience.
AI Agents 3 - Agentic Design Patterns
This video from Prof. Ghassemi Lectures and Tutorials provides an excellent breakdown of these three components and illustrates the perception-planning-action loop with two clear examples: ChatGPT and a hypothetical health AI agent.
Watch from 00:35 to 05:57. Focus on how the agent (like ChatGPT) perceives user input, uses its LLM 'brain' for planning, and leverages tools (like web search) and memory to formulate an action (the response).
This model of an LLM orchestrating tools and memory is central to all modern agentic frameworks.
Building an AI Coding Agent from Scratch Using Python
For a developer-focused perspective, let's look at the article 'Building an AI Coding Agent from Scratch Using Python'.
Read the sections 'What Is an AI Coding Agent?' and 'Core Idea Behind an Agent'. This reinforces the idea that an agent is fundamentally a loop combining an LLM and a set of tools.
4. Designing a Basic Agent: A Walkthrough
Let's solidify these concepts by designing a simple agent. Our goal is to create a "Weather Bot" that can answer questions about the current weather in a given city.
Here's how we'd design it using the perception-action loop:
- Goal: To provide the current weather for a specified city.
- Environment: A command-line interface where a user can type questions.
- Tool: We need a way to get weather data. We'll define a single tool, a Python function
get_weather(city: str) -> str, which internally calls a weather API.
Now, let's trace one cycle of the loop.
1. Perception
The agent perceives an input from the environment.
- User Input (Percept):
"What is the weather like in Tokyo?"
2. Planning / Thinking
The user's input is passed to the LLM, along with a system prompt that defines its role, its available tools, and how it should respond.
- System Prompt:
You are a helpful assistant. Your goal is to answer questions about the weather. You have access to the following tool: - get_weather(city: str): Returns the current weather for the given city. To use a tool, respond with a JSON object like: {"tool_name": "get_weather", "arguments": {"city": "Tokyo"}} If you don't need a tool, respond directly to the user. - LLM Reasoning: The LLM analyzes the user's query
"What is the weather like in Tokyo?". Guided by the system prompt, it recognizes the intent to find weather and identifies "Tokyo" as the city. It concludes that it needs to use theget_weathertool. - LLM Output (Plan): The LLM generates the structured action plan.
{"tool_name": "get_weather", "arguments": {"city": "Tokyo"}}
3. Action & Observation
An orchestrator program outside the LLM parses this JSON output.
- Action Execution: The orchestrator sees the request to use the
get_weathertool and calls the corresponding Python function:get_weather("Tokyo"). - Observation (New Percept): The function executes, calls the external weather API, and returns a result, for example:
"15°C and sunny".
4. Planning / Thinking (Round 2)
The loop continues. The result from the tool is fed back to the LLM as new context.
- New Context for LLM:
You previously decided to call the tool 'get_weather' with the city 'Tokyo'. The tool returned the following result: "15°C and sunny". Now, formulate a natural language response for the user. - LLM Reasoning: The LLM synthesizes this information into a user-friendly sentence.
- LLM Output (Final Action):
"The current weather in Tokyo is 15°C and sunny."
This final text is then displayed to the user in the command-line interface, completing the loop for this query.
This iterative perceive -> plan -> act -> observe -> repeat cycle is the essence of agentic behavior. It allows the system to break down a problem, gather information, and then formulate a final response.
The Agentic AI Handbook: A Beginner's Guide to ...
The 'Agentic AI Handbook' provides an excellent pseudocode example that formalizes this loop.
Read the sections 'Planning' and 'Code Snippet and Real-World Examples'. The 'Planning' section contains a simplified loop perceive -> plan -> act -> learn. The 'Code Snippet' section shows how this could be structured within a Python class, which should be very intuitive for you.
Test your understanding!
Imagine you are designing an agent to help you manage a music playlist. Your goal is to add a specific song to a specific playlist.
-
Tools available:
find_song_id(song_name: str, artist_name: str) -> stradd_song_to_playlist(song_id: str, playlist_name: str) -> str
-
User Input:
"Add 'Yesterday' by The Beatles to my 'Classics' playlist."
Outline the steps of the perception-action loop for this task. What would the LLM plan in the first step? What would the second step be after observing the result of the first?
Show answer
- Perception: Agent perceives the user request:
"Add 'Yesterday' by The Beatles to my 'Classics' playlist." - Planning (Step 1): The LLM recognizes it needs a
song_idbefore it can add the song. It plans to call thefind_song_idtool.- LLM Output:
{"tool_name": "find_song_id", "arguments": {"song_name": "Yesterday", "artist_name": "The Beatles"}}
- LLM Output:
- Action & Observation (Step 1): The system executes
find_song_id("Yesterday", "The Beatles")and gets a result.- Tool Result:
"7iN1s7xHE4ifF5povM6A48"(a fictional song ID).
- Tool Result:
- Planning (Step 2): This result is fed back to the LLM. It now has the
song_idand knows the target playlist is 'Classics'. It plans the final action.- LLM Output:
{"tool_name": "add_song_to_playlist", "arguments": {"song_id": "7iN1s7xHE4ifF5povM6A48", "playlist_name": "Classics"}}
- LLM Output:
- Action & Observation (Step 2): The system executes the function.
- Tool Result:
"Successfully added song to 'Classics' playlist."
- Tool Result:
- Planning (Step 3): This final confirmation is sent to the LLM, which then formulates a response to the user.
- LLM Output:
"Done! I've added 'Yesterday' by The Beatles to your 'Classics' playlist."
- LLM Output:
Conclusion
In this lesson, we established the fundamental principles of agentic AI. You've learned that an agent is more than just a reactive model; it's an autonomous system that operates within an environment to achieve a goal.
Key Takeaways:
- An AI agent perceives its environment and acts upon it to achieve goals.
- The core operating principle of any agent is the perception-action loop.
- Modern LLM-based agents consist of three main components: an LLM (the brain), a set of tools (the hands), and memory.
- The agent's "thinking" process involves the LLM planning a sequence of actions, often involving tool calls, to break down a complex task into manageable steps.
This perception-action loop is the foundational pattern we will build upon throughout this module. By mastering this design, you'll be able to create sophisticated AI systems capable of complex problem-solving.
Preview of the Next Lesson:
We've designed a basic agent conceptually. In the next lesson, we will get practical and start building. Our objective will be to implement function calling (tool use) to allow an LLM to interact with external APIs. This will be our first step in turning the designs we discussed today into working code.