Hello! Welcome to the fifth lesson in our module on Reinforcement Learning Foundations.
In our last lesson, we explored how to evaluate a given policy without a model of the environment, using Monte Carlo (MC) and Temporal Difference (TD) prediction. We learned to estimate the state-value function, , purely from experience. But evaluation is only half the battle. Our ultimate goal is to find the best policy.
Today, we make the crucial leap from prediction to control. We'll learn how to not just evaluate, but to actively improve our policy and find the optimal way to act. This lesson directly addresses the learning outcome: Implement model-free control algorithms: SARSA and Q-Learning. These two algorithms are foundational to reinforcement learning and serve as the perfect entry point into the critical concepts of on-policy and off-policy learning.
1. From Prediction to Control: The Action-Value Function
In our previous model-based lessons (Value and Policy Iteration), we could improve a policy by looking one step ahead. With a state-value function , we'd check all actions, see where they lead using the model , and pick the action leading to the best outcome.
In a model-free setting, this is impossible. We don't have the transition probabilities . So, how do we choose the best action?
The answer is to shift from learning state values to learning action-values. We introduce the action-value function, , which represents the expected return from starting in state , taking action , and then following policy thereafter.
If we have the optimal action-value function, , finding the optimal policy becomes trivial—no model needed! We simply act greedily at each state:
Our goal is now to estimate using the TD learning principles from our last lesson.
To set the stage, let's first get a clear understanding of this "quality" function and why model-free methods are necessary.
Q-Learning: Model Free Reinforcement Learning and Temporal Difference Learning
Professor Steve Brunton provides an excellent motivation for model-free learning and re-introduces the quality function, Q, which is central to today's lesson.
Watch from 00:08 to 06:03. Focus on: The definition of the quality function, Q(s,a). The key reason we need model-free algorithms: we often don't know the system dynamics (P) or reward function (R).
2. Q-Learning: Off-Policy TD Control
Q-learning is arguably the most famous RL algorithm. It's a model-free, TD control method that directly approximates the optimal action-value function, , regardless of the policy being followed.
The update rule for Q-learning looks very similar to the TD(0) update we saw for V-functions, but applied to Q-values:
Let's dissect this:
- We take an action in state and observe the reward and next state .
- The crucial part is the TD target: . To estimate the value from the next state, Q-learning uses the maximum Q-value possible from . It assumes we will take the best possible action from that point onwards.
- This makes Q-learning an off-policy algorithm. The policy it learns about (the optimal, greedy policy) is different from the policy it uses to generate actions (the behavior policy, which must explore, e.g., -greedy). It learns about being greedy, even while it's behaving non-greedily.
This diagram clearly shows the update mechanism.

For a detailed walkthrough of the Q-learning algorithm and its off-policy nature, watch the following segment.
Q-Learning: Model Free Reinforcement Learning and Temporal Difference Learning
Steve Brunton explains the Q-learning update rule and the concept of off-policy learning, highlighting its major benefits.
Watch from 22:54 to 26:16. Pay close attention to how the max operator in the TD target enables off-policy learning and allows the agent to learn from suboptimal actions, replays, or even imitation.
Implementing Q-Learning
Given your background in Python, let's look at a practical implementation. We'll use the FrozenLake-v1 environment from the Gymnasium library—a classic grid world where an agent must navigate a frozen lake to a goal, avoiding holes.
Q-Learning Tutorial 1: Train Gymnasium FrozenLake-v1 with Python Reinforcement Learning
This tutorial by Johnny Code provides a step-by-step guide to implementing Q-learning from scratch in Python to solve the FrozenLake environment.
Watch from the beginning to 05:53. This will guide you through: Setting up the Gymnasium environment. Initializing the Q-table (a 2D NumPy array for states × actions). Implementing the epsilon-greedy policy for exploration. Writing the main training loop that applies the Q-learning update rule.
As you watch, consider how you would structure this code. The core components are:
- A loop over episodes.
- An inner loop for steps within an episode.
- An action selection mechanism (-greedy).
- Executing the action and getting
(next_state, reward, done). - Updating the Q-table using the Q-learning rule.
For a formal pseudocode representation and a well-structured Python implementation, you can refer to the resource "Temporal difference reinforcement learning". It presents the algorithm clearly in Algorithm 4.
3. SARSA: On-Policy TD Control
SARSA is the on-policy sibling of Q-learning. Its name is a mnemonic for the sequence of events involved in its update: State, Action, Reward, State, Action.
The SARSA update rule is subtly but profoundly different from Q-learning:
Notice the TD target: .
- Instead of taking the maximum Q-value at the next state, SARSA uses the Q-value of the actual next action, , that the agent chose to take according to its behavior policy (e.g., -greedy).
- This makes SARSA an on-policy algorithm. It learns the value of the exact policy it is following, including its exploratory actions. It doesn't make optimistic assumptions about acting greedily in the future.
The flowcharts below provide a great side-by-side comparison of the procedural differences.

Let's get a clear explanation of SARSA and its on-policy nature.
Q-Learning: Model Free Reinforcement Learning and Temporal Difference Learning
Here, Steve Brunton explains the SARSA algorithm, directly contrasting it with Q-learning.
Watch from 26:16 to 29:17. Focus on how the SARSA update rule differs and why this requires the agent to be 'on-policy' for the estimate to be correct.
Implementing SARSA is very similar to implementing Q-learning. The main difference is that you must choose the next action, , before you can calculate the TD target and update the Q-value for the current state-action pair, . The gibberblot.github.io resource provides excellent pseudocode in Algorithm 5 and a Python implementation that highlights this difference.
4. On-Policy vs. Off-Policy: The Cliff Walking Example
So, which is better? The answer depends on the problem. The classic "Cliff Walking" environment provides the perfect illustration of their different behaviors.
The Setup: An agent starts at one corner of a grid and must reach the goal at the opposite corner. There's a "cliff" along the bottom edge.
- Optimal Path: Move right along the edge of the cliff. This is the shortest path but is very risky. One wrong (exploratory) move, and the agent falls off, receiving a large negative reward.
- Safe Path: Move up and around the cliff. This path is longer (and thus accumulates more small negative rewards per step) but is safe from falling.
Temporal difference reinforcement learning
This article provides a fantastic, detailed walkthrough of the Cliff World example, comparing the policies learned by Q-learning and SARSA.
Read the section 'SARSA vs. Q-learning example: Cliff world'. Pay close attention to the policies visualized for both algorithms and the explanation for why they differ.
Here's the takeaway from the Cliff Walking example:
- Q-Learning (Off-Policy): Learns the optimal policy (the risky path along the cliff). Because its update rule uses , it ignores the consequences of the exploratory actions that lead to falling off the cliff. It's an optimistic learner.
- SARSA (On-Policy): Learns the safe policy (the longer path). Because its update rule uses the Q-value of the next actual action, , the large negative rewards from exploratory "falls" are incorporated into the Q-values of states near the cliff. It learns that being near the cliff is dangerous and finds a safer, albeit suboptimal, path.
This reveals a fundamental trade-off: optimality vs. safety during training.
Test your understanding!
Imagine you are developing an RL agent for two different tasks:
- A simulated agent learning the absolute best strategy for a stock trading game, where you can run millions of episodes at no real cost.
- A physical robot learning to navigate a factory floor with expensive equipment. Mistakes during training could be very costly.
Which algorithm, Q-learning or SARSA, would you choose as a starting point for each task, and why?
Show answer
-
Stock Trading Game: Q-learning is the better choice. The goal is to find the absolute optimal policy, and the cost of "bad" exploratory actions during training is zero (it's just a simulation). Q-learning's off-policy nature allows it to learn the optimal path without being overly conservative due to exploration.
-
Physical Robot: SARSA is the safer and more appropriate starting point. The on-policy nature of SARSA means it will be sensitive to the real-world consequences of its exploratory actions. It will learn a "safer" policy that avoids costly mistakes, which is paramount when physical damage or safety is a concern. The final policy might be slightly suboptimal, but the training process will be much less hazardous.
This detailed comparison video will solidify the practical differences between the two algorithms.
Q-Learning: Model Free Reinforcement Learning and Temporal Difference Learning
Finally, this segment provides a direct, side-by-side comparison of Q-learning and SARSA, summarizing their strengths and weaknesses.
Watch from 29:17 to 34:48. This is a crucial summary of the entire lesson. Focus on the trade-offs regarding learning speed, safety, and cumulative reward during training.
Conclusion
In this lesson, we moved from model-free prediction to model-free control, learning how to find optimal policies directly from experience. We implemented and contrasted two of the most important algorithms in the RL canon.
Key Takeaways:
- Model-Free Control uses the action-value function to select the best actions without needing a model of the environment.
- Q-Learning is an off-policy algorithm. It learns the optimal (greedy) policy by using the operator in its update, making it "optimistic" and often faster at finding the best possible path.
- SARSA is an on-policy algorithm. It learns the value of the policy it is actually following (including exploration), making it more "realistic" and sensitive to the costs of mistakes during training.
- The choice between them represents a fundamental optimality vs. safety trade-off. SARSA is often preferred when training mistakes are costly, while Q-learning is preferred when the goal is to find the best possible policy in a safe (e.g., simulated) environment.
Preview of the Next Lesson:
We've just had a very practical introduction to the concepts of on-policy and off-policy learning. In our next lesson, we will formalize this distinction. We will explore more deeply what it means for an algorithm to be on-policy or off-policy and discuss the broader implications and applications of each approach, setting the stage for more advanced algorithms.