Skip to main content
Create your own

PPO with Clipped Objective

Hello! Welcome to our next lesson on Deep Reinforcement Learning.

In our previous session, we built an Advantage Actor-Critic (A2C) agent. We saw how using a critic to provide an advantage estimate, , creates a more stable learning signal than the noisy Monte Carlo returns used in REINFORCE. However, A2C still has a critical vulnerability: a bad batch of data can lead to a large, destructive policy update, causing performance to collapse without recovery. We called this an "unconstrained" update.

Today, we will address this problem by implementing Proximal Policy Optimization (PPO). PPO is one of the most successful and widely-used RL algorithms, largely because it solves this stability issue with a simple yet effective mechanism. Your learning outcome for this lesson is to implement Proximal Policy Optimization (PPO) with clipping for stable training.

1. The Core Idea: Constraining Policy Updates

The central innovation of PPO is to ensure that the updated policy does not stray too far from the old one in a single step. It accomplishes this by modifying the objective function.

Instead of just multiplying the advantage by the log-probability of an action, PPO looks at the ratio of the probabilities between the new and old policies:

  • If , the new policy makes the action more likely.
  • If , the new policy makes the action less likely.

PPO's objective is to increase the probability of actions with positive advantages and decrease it for those with negative advantages, but it "clips" the objective to discourage the ratio from becoming too large or too small. This acts as a set of guardrails, preventing the policy from changing too dramatically.

Let's begin with a conceptual overview that explains this core mechanism in a clear, intuitive way.

Simply Explaining Proximal Policy Optimization (PPO) | Deep Reinforcement Learning

This video from Johnny Code provides an excellent, simplified explanation of the PPO algorithm, focusing on the intuition behind each step.

Please watch from 20:04 to 27:07. This section masterfully breaks down the PPO clip objective function. Pay close attention to: The explanation of the probability ratio, r. How the clip function limits this ratio within [1-ε, 1+ε]. The graphs illustrating how the final objective changes for positive and negative advantages. The contrast with the older REINFORCE algorithm to highlight why this clipping is so important for stability.

2. The PPO Clipped Surrogate Objective

As you saw in the video, the magic of PPO lies in its unique objective function. Let's look at the formal definition, which is known as the clipped surrogate objective:

Here, is the advantage estimate at time , and (epsilon) is a small hyperparameter (e.g., 0.2) that defines the clipping range.

This formula might seem complex, but its behavior is straightforward and can be broken down into two main cases based on the sign of the advantage.

Proximal Policy Optimization — Key Equations

The OpenAI Spinning Up documentation provides the most concise and accurate explanation of this objective function. Let's read through it.

Read the section titled 'Key Equations'. It starts with the formula for L_CLIP and provides a brilliant, intuitive breakdown for the cases when the 'Advantage is positive' and 'Advantage is negative'. Focus on understanding why 'the new policy does not benefit by going far away from the old policy' in both scenarios.

To summarize the logic from the reading, the min function in the objective has a clear purpose:

  • If (good action): The objective is clipped from above by . This encourages making the action more likely (increasing ) but removes the incentive for making it too much more likely.
  • If (bad action): The objective is clipped from below by . This encourages making the action less likely (decreasing ) but removes the incentive for making it too much less likely.

This visual from Hugging Face perfectly captures all possible scenarios:

PPO Clipping Mechanism Visualization
This diagram illustrates the PPO clipping mechanism. The table and graphs show how the objective function is bounded for both positive and negative advantages, preventing extreme policy updates.
Test your understanding!

Suppose and we have an action with a positive advantage . What would be the contribution to the objective function, , if the probability ratio were:

  1. (a small, good update)
  2. (a large, risky update)
Show answer

The clipping range is .

  1. For :

    • Unclipped term: .
    • Clipped term: .
    • The objective is . The update is not clipped.
  2. For :

    • Unclipped term: .
    • Clipped term: .
    • The objective is . The update is clipped. Even though the policy proposed a large change that would have yielded a "score" of 15, its contribution is limited to 12. This removes the incentive for making such large updates.

3. PPO Implementation from Scratch

Now that we have a solid grasp of the theory, let's dive into the implementation. We will build a complete PPO agent in PyTorch. As PPO is an Actor-Critic algorithm, the overall structure will be familiar from our last lesson, but with a few key differences:

  1. Memory Management: PPO is an on-policy algorithm that learns from fixed-length batches of experience. We need a memory structure to collect these trajectories and serve them up for training in mini-batches.
  2. Advantage Calculation: We'll use Generalized Advantage Estimation (GAE), a more advanced technique that balances bias and variance to get a better advantage estimate than the simple one-step TD error. The formula is:

    where is the TD error, is the discount factor, and is a smoothing parameter (e.g., 0.95).
  3. The Learning Function: This is where we'll implement the PPO clipped surrogate objective for the actor loss and a standard MSE loss for the critic.

The following video is a comprehensive, line-by-line tutorial that will guide us through the entire process. Given your background in software engineering, walking through a full build is the most effective way to master the algorithm.

PPO Actor-Critic Architecture Diagram
A high-level view of the data flow in a PPO agent. Experience from the environment is stored in a trajectory memory. Mini-batches are then sampled to update both the Actor and Critic networks based on the PPO objective and value loss, respectively.

Part 1: Theory Recap and Memory Implementation

Let's start with the video's introduction to PPO and the implementation of the memory buffer.

Proximal Policy Optimization (PPO) is Easy With PyTorch | Full PPO Tutorial

We'll use the 'Machine Learning with Phil' tutorial as our primary guide. This first segment covers the PPO concepts and the specialized memory class we need.

Watch from the beginning to 23:59. 00:00 - 02:00: Motivation for PPO and a high-level overview of the clipped loss. 02:00 - 03:44: Explanation of using mini-batch SGD over fixed-length trajectories. 11:11 - 13:06: A clear explanation of Generalized Advantage Estimation (GAE) and the critic's loss function. 19:08 - 23:59: The Python implementation of the PPOMemory class. Pay attention to how store_memory collects data and how generate_batches shuffles and yields mini-batches for training.

Part 2: Actor and Critic Networks

Next, we'll define the neural network architectures for our Actor and Critic. The video implements them as two separate networks, which is a clear and effective design.

Proximal Policy Optimization (PPO) is Easy With PyTorch | Full PPO Tutorial

Let's continue with the implementation of the Actor and Critic networks.

Watch from 23:59 to 32:49. ActorNetwork (23:59 - 27:40): Notice the use of a softmax activation to output a probability distribution, which is then wrapped in a Categorical distribution object from PyTorch. This handles action sampling and log-probability calculations for us. CriticNetwork (27:40 - 32:49): Observe that the critic's output is a single linear node, representing the state-value V(s).

Part 3: The Agent and the Learning Logic

This is the most critical part of the implementation. We will create the Agent class that brings everything together and contains the core learn() method.

Proximal Policy Optimization (PPO) is Easy With PyTorch | Full PPO Tutorial

Now for the main event: the Agent class and its learn() method, where the PPO objective is implemented.

Watch from 35:37 to 48:55. choose_action (35:37 - 38:46): See how the agent gets the action, its log probability, and the critic's value for a given state. These are all stored in memory. learn() (38:46 - 48:55): This is the heart of the algorithm. Follow the implementation step-by-step: GAE Calculation: A nested loop implements the GAE formula to compute the advantage for each timestep. Mini-batch Loop: The code iterates through the mini-batches provided by our memory class. Ratio Calculation: It gets new probabilities from the current actor network, computes the ratio r_t = \exp(\log\pi_{new} - \log\pi_{old}). Actor Loss: It calculates weighted_probs (the unclipped term) and weighted_clipped_probs (the clipped term), then takes min() of the two and the negative mean (since optimizers minimize). Critic Loss: It calculates the returns (\hat{A}_t + V_{old}) and computes the mean squared error against the new critic's prediction. Update: It sums the losses, backpropagates, and steps the optimizers.

Part 4: Training Loop and Results

Finally, let's see how the agent is trained and how it performs on the CartPole environment.

Proximal Policy Optimization (PPO) is Easy With PyTorch | Full PPO Tutorial

Let's wrap up by looking at the main training loop that puts the agent to work.

Watch from 48:55 to the end. This shows how the agent interacts with the environment, stores memories, and periodically calls the learn() function. Observe the resulting learning curve.

4. Hyperparameter Tuning and Common Pitfalls

While PPO is known for its robustness, its performance is still sensitive to hyperparameter choices. Getting an implementation to work well often involves careful tuning. The resource below provides an excellent cheat sheet for common issues and how to fix them.

PPO Hyperparameter tuning and common pitfalls

This table from a Digital Ocean tutorial is an invaluable resource for practical PPO implementation. It summarizes common pitfalls and recommended heuristics for key hyperparameters.

Review the table in this section. Pay particular attention to the advice on Learning Rate, Clip Range (ε), Batch Size, and Advantage Normalization. This is practical knowledge that will save you a lot of time when training your own agents.

Conclusion

Congratulations! You have now implemented Proximal Policy Optimization, one of the most important and effective algorithms in modern Deep Reinforcement Learning. You've gone from the basic idea of policy gradients to the stable, powerful, and constrained updates that make PPO so reliable.

Key Takeaways:

  • PPO's Goal: To improve the policy without taking destructively large steps, ensuring stable learning.
  • Core Mechanism: It uses a clipped surrogate objective that limits how much the probability ratio can influence the update.
  • The Objective: . This acts as a pessimistic lower bound on the policy improvement.
  • Implementation: PPO is an on-policy, actor-critic algorithm that learns from fixed-length trajectories, typically uses GAE for advantage estimation, and performs multiple mini-batch updates per data collection phase.

Preview of the Next Lesson:

We have now covered the three foundational deep RL algorithms: REINFORCE, A2C, and PPO. These methods form the bedrock of many advanced applications.

Starting with our next module, we will shift our focus to the world of Large Language Models (LLMs). You might be surprised to learn that techniques inspired by reinforcement learning are crucial for making these models safe and helpful. We will begin by exploring how to align LLMs with human instructions through a process called Supervised Fine-Tuning (SFT).

Can't find a good explanation? Sign up and we'll make it for you

Sign up