Hello! Welcome to the final lesson in our module on Reinforcement Learning Foundations.
In our last lesson, we implemented SARSA and Q-learning and saw how they produced different behaviors in the "Cliff Walking" environment. We identified SARSA as "on-policy" and Q-learning as "off-policy" based on their update rules. Today, we're going to formalize that distinction.
This lesson addresses the learning outcome: Explain the difference between on-policy and off-policy learning. We will move beyond specific algorithms to understand the general principles, which is a crucial conceptual step before we advance to more complex Deep Reinforcement Learning methods.
A Quick Recap: Where We Are
Reinforcement learning algorithms can be broadly categorized as model-based or model-free. We are currently focused on model-free methods, which learn directly from experience without a model of the environment's dynamics. Within model-free learning, the on-policy vs. off-policy distinction is one of the most important organizational concepts.

1. The Core Distinction: Behavior vs. Target Policies
The difference between on-policy and off-policy learning boils down to the roles a policy can play. In model-free control, a policy is doing two jobs simultaneously:
- It's used to make decisions and generate experience data by interacting with the environment.
- It's the thing we are trying to improve to make it better.
This leads to two key definitions:
- Behavior Policy (): The policy the agent actually uses to choose actions and explore the environment. Think of this as the "data-gathering" policy.
- Target Policy (): The policy the agent is trying to learn and improve. This is the policy we want to eventually become optimal.
With these definitions, the distinction becomes clear:
- On-Policy methods learn about the value of the policy they are currently executing. The behavior policy and the target policy are the same ().
- Off-Policy methods learn about the value of a policy that is different from the one they are using to generate actions. The behavior policy and the target policy are different ().
This video provides an excellent and clear explanation of this fundamental concept.
Monte Carlo And Off-Policy Methods | Reinforcement Learning Part 3
This video from the Mutual Information channel introduces the concepts of behavior and target policies, which are the key to understanding the on-policy vs. off-policy distinction.
Watch from 21:46 to 23:01. Focus on how the two roles of a policy—generating data and being improved—are separated into the 'behavior policy' and the 'target policy'.
2. SARSA and Q-Learning Through the Policy Lens
Let's revisit our two main algorithms from the last lesson and see how they fit this formal definition.
SARSA: On-Policy
The SARSA update rule is:
The crucial term is . To perform this update, the agent must first select the next action, , using its current behavior policy (e.g., -greedy). The policy used to select for the update is the same as the policy being evaluated and improved. Thus, SARSA is on-policy.
Q-Learning: Off-Policy
The Q-learning update rule is:
Here, the update uses . This means the target policy () is a greedy policy that always chooses the action with the highest Q-value. However, the behavior policy () used to select the action that generates the experience is typically -greedy to ensure exploration. Since the greedy target policy is different from the -greedy behavior policy, Q-learning is off-policy.
This image visually captures the difference in their update targets. SARSA considers the expected value under its policy (which includes exploratory moves), while Q-learning greedily takes the maximum value.

For a more detailed textual explanation with pseudocode, the following article is highly recommended.
On-Policy vs Off-Policy Reinforcement Learning
This article from Georgia Tech's CORE Robotics Lab clearly defines the behavior and update (target) policies and walks through SARSA and Q-learning as concrete examples.
Read sections 1 ('What is on-policy versus off-policy?') and 2 ('Examples'). Pay attention to how the article defines the behavior and update policies and highlights the key lines of code that make SARSA on-policy and Q-learning off-policy.
3. Why It Matters: The Profound Implications
This distinction is not just academic; it has massive practical consequences for how we design and use RL agents.
Data Efficiency and Experience Replay
This is arguably the most significant implication.
- On-policy methods are sample-inefficient. When the policy is updated, the experiences generated by the old policy are no longer "on-policy." They must be thrown away. The agent can only learn from the data generated by its current, exact policy.
- Off-policy methods are sample-efficient. Because the behavior policy can be different from the target policy, an off-policy agent can learn from experiences generated by any policy. This enables experience replay, where experiences are stored in a large memory buffer. The agent can then sample from this buffer repeatedly to learn, making much more efficient use of each interaction with the environment.
This ability to reuse past data is a cornerstone of modern deep reinforcement learning.
Exploration vs. Exploitation
- In on-policy learning, the exploration strategy is part of the policy being learned. The agent must "live with" the consequences of its exploration. This is why SARSA learned the safer, longer path in the Cliff Walking problem—it had to account for the possibility of an -random action leading it off the cliff.
- In off-policy learning, exploration and control are decoupled. The agent can use a highly exploratory behavior policy (e.g., very random) to gather a wide range of data while the target policy being learned remains purely greedy and optimal.
Learning from Others
Off-policy learning allows an agent to learn from data it didn't generate itself. For example, an agent can learn by:
- Observing a human expert (imitation learning).
- Using large, pre-existing datasets of trajectories.
- Multiple agents sharing their experiences to learn a single policy.
This is impossible for a pure on-policy method, which can only learn from its own actions.
The following video discusses the motivation for off-policy methods and touches on a key technical aspect.
Monte Carlo And Off-Policy Methods | Reinforcement Learning Part 3
Let's return to the Mutual Information video to understand the practical motivation for off-policy methods and get a glimpse of the theory behind them.
Watch from 23:01 to 24:31. Note the key motivation: the ability to use pre-existing data. Also, listen for the mention of 'importance sampling,' which is the mathematical technique used to correct for the fact that the behavior and target distributions are different.
Test your understanding!
An RL agent is learning to play chess. It stores every game it has ever played in a large database. After every new game, it samples 1,000 random state-action-reward transitions from this entire database to update its Q-function.
Is this approach fundamentally on-policy or off-policy? Why?
Show answer
This is a classic off-policy approach.
The agent is learning from a "replay buffer" of past experiences. The policy that generated the games at the beginning of training is very different from the much-improved policy that exists after thousands of games. Because the agent is learning from data generated by policies other than its current one, it is, by definition, off-policy.
An on-policy agent would have to discard the data from each game after updating its policy once, as that data would be "off-policy" for the new, updated policy.
4. Summary of Trade-offs
Let's conclude with a summary of the trade-offs, which we first saw in the last lesson but can now understand through the formal on/off-policy lens.
On-Policy vs Off-Policy Reinforcement Learning
The Georgia Tech article provides an excellent summary of the practical implications and trade-offs.
Read section 3 ('Implications'). This section provides a concise breakdown of how the different learning strategies lead to different policy outcomes and when you might choose one over the other.
| Aspect | On-Policy (e.g., SARSA) | Off-Policy (e.g., Q-Learning) |
|---|---|---|
| Policy Learned | Learns the value of the current (often -soft) policy. | Learns the value of the optimal (greedy) policy. |
| Data Usage | Sample-inefficient. Must discard old data. | Sample-efficient. Can use experience replay. |
| Exploration | Exploration is coupled with the target policy, leading to more "conservative" or "safer" behavior. | Exploration is decoupled, allowing for more aggressive exploration while still learning an optimal policy. |
| Variance | Generally lower variance updates, leading to more stable convergence. | Can have higher variance due to importance sampling correction, but techniques exist to manage this. |
| Use Cases | Good for tasks where safety during training is critical (e.g., robotics), as it accounts for exploration mistakes. | Excellent for simulated environments and tasks where data is expensive, as it can reuse data to find the true optimal policy. |
Conclusion
In this lesson, we formalized the crucial distinction between on-policy and off-policy learning. This is one of the most important concepts in reinforcement learning.
Key Takeaways:
- The difference hinges on the behavior policy (used for data collection) and the target policy (the policy being improved).
- In on-policy methods, the behavior and target policies are the same. SARSA is the classic example.
- In off-policy methods, the behavior and target policies are different. Q-learning is the classic example.
- This difference has profound implications for data efficiency, exploration, and the ability to learn from diverse data sources.
- Off-policy learning's ability to use experience replay makes it far more sample-efficient and is a foundational concept for many advanced algorithms.
Preview of the Next Lesson:
With a solid understanding of Markov Decision Processes, Bellman equations, and the on/off-policy distinction, we are now ready to leave the world of tabular methods behind. In the next module, we will begin our study of Deep Reinforcement Learning. Our first topic will be the Deep Q-Network (DQN), an algorithm that combines Q-learning with deep neural networks and leverages experience replay to achieve superhuman performance in Atari games.