Skip to main content
Create your own

Actor-Critic Architectures: A2C Implementation

Hello! Let's dive into the next stage of our reinforcement learning journey.

In our last lesson, we implemented REINFORCE, our first policy gradient algorithm. We saw that while it works by directly optimizing the policy, it suffers from high variance because it relies on the noisy, full-episode Monte Carlo return as its learning signal. A single unlucky outcome could unfairly penalize a whole sequence of good actions. We concluded by introducing the idea of subtracting a baseline, like the state-value function , to create a more stable learning signal: the advantage.

This brings us directly to today's topic. Instead of just assuming we have a value function, we are going to learn it alongside the policy. This powerful combination is the core of Actor-Critic methods. Your goal for this lesson is to build an Actor-Critic architecture, specifically the popular variant known as Advantage Actor-Critic (A2C).

1. The Actor-Critic Framework

The name "Actor-Critic" perfectly describes the architecture. We have two distinct components, typically implemented as two neural networks or two "heads" of a single network:

  • The Actor is the policy, . It observes the state and decides which action to take. This is the component we had in REINFORCE.
  • The Critic is a value function, typically the state-value function . It doesn't choose actions; it only evaluates the states visited by the Actor, judging "how good" they are.

The two work in a feedback loop: the Actor performs, and the Critic provides feedback, which the Actor uses to improve.

Actor-Critic Architecture Diagram
A block diagram illustrating the fundamental components of an Actor-Critic architecture. The Actor selects an action, the environment provides the next state and a reward, and the Critic uses this information to provide a learning signal (the TD-error) back to the Actor.

To get a more intuitive feel for this dynamic, let's start with a short conceptual overview.

Advantage Actor Critic (A2C)

This article from the Hugging Face blog uses a great analogy to explain the core concept of the Actor-Critic relationship.

Read the first two sections, 'Reducing variance with Actor-Critic methods' and 'The Actor-Critic Process'. Focus on the analogy of the player (Actor) and the friend giving feedback (Critic) and how they improve together.

2. From Policy Gradients to Advantage Updates

Recall the REINFORCE update rule with a baseline:

The term is an estimate of the advantage function .

Actor-Critic methods refine this in two key ways:

  1. We explicitly learn the value function with a parameterized Critic network.
  2. Instead of waiting for the full Monte Carlo return , we use a one-step Temporal Difference (TD) target, .

This gives us the TD error, , which serves as a low-variance, single-step estimate of the advantage:

This becomes the learning signal for both the Actor and the Critic.

  • Critic's Goal: To make its value estimates more accurate. It is trained to minimize the TD error. A common loss function is the mean squared error of the TD error:
  • Actor's Goal: To take actions that lead to positive advantage. Its loss is the negative log-probability of the action, scaled by the advantage. We use .detach() on the advantage term because we only want the Critic to provide a numeric score, not to receive gradients from the Actor's loss.

This structure is the foundation of the Advantage Actor-Critic (A2C) algorithm.

Advantage Actor-Critic (A2C) Architecture Diagram
This diagram shows the information flow in A2C. The Actor and Critic (Policy and Value networks) take the current state. The Actor produces an action. After the environment step, the Critic uses the reward and new state to compute the advantage, which is used as the learning signal for the Actor.

3. Building an A2C Agent from Scratch

Now, let's translate this theory into a working implementation. We will build an A2C agent step-by-step. Since you're proficient in Python, walking through a detailed coding tutorial is the most effective way to master this.

The following video provides an exceptionally thorough guide to building an Actor-Critic agent using TensorFlow 2. It covers the theory, the network architecture, the agent logic, and the final training loop. We will work through it in sections.

Everything You Need To Master Actor Critic Methods | Tensorflow 2 Tutorial

The 'Machine Learning with Phil' channel provides a detailed, from-the-ground-up implementation of a basic Actor-Critic agent. We'll use this as our primary guide.

Watch from the beginning to 12:33. This part covers the core theoretical foundations of the algorithm that will be implemented. 00:42 - 09:08: This serves as a solid refresher on MDPs, policies, returns, and value functions. Pay attention to how the return G_t is defined recursively. 09:08 - 12:33: This is the most critical section. It explains how the Actor and Critic networks work together, the concept of Temporal Difference (TD) learning, and defines the crucial delta term (our TD error / advantage) and the loss functions for both the Actor and the Critic.

Test your understanding!

In the video at timestamp 11:55, the delta term is defined as:

And the loss functions are given for the Critic () and the Actor ().

Why do we want to make the Actor's loss negative? What would happen if we tried to minimize ?

Show answer

Gradient descent minimizes a loss function. Our goal is to make actions that lead to a positive advantage () more likely.

The term is the direction in parameter space that increases the probability of taking action in state .

  • If (a good action), we want to move our parameters in the direction of . Minimizing is equivalent to performing gradient ascent on , which achieves this.
  • If we tried to minimize , the optimizer would push the parameters in the opposite direction of when . This would mean punishing good actions and rewarding bad ones, preventing the agent from learning.

Network and Agent Implementation

With the theory in place, let's implement the model. A common and efficient design is to use a single neural network with a shared "body" for feature extraction and two separate "heads":

  1. A policy head that outputs logits for the action distribution (the Actor).
  2. A value head that outputs a single number representing the state value (the Critic).

This allows the Critic to benefit from the representations learned for the Actor, and vice-versa.

Everything You Need To Master Actor Critic Methods | Tensorflow 2 Tutorial

Let's continue with the implementation part of the video. This section details the network architecture and the core agent logic.

Watch from 12:33 to 31:06. This is the heart of the implementation. 12:33 - 19:42 (The Algorithm and Network): First, the video outlines the high-level algorithm. Then, it implements the ActorCriticNetwork class. Note how it has a shared fc1 and fc2 layers, but two distinct output layers: v (for the critic) and pi (for the actor). 19:42 - 31:06 (The Agent): This section implements the Agent class. Pay close attention to: choose_action(): How tensorflow_probability is used to create a Categorical distribution from the actor's output probabilities to sample an action. learn(): This is the core update logic. Follow how it uses tf.GradientTape to compute the state_value and new_state_value, calculate delta (the advantage), define the actor_loss and critic_loss, and finally compute and apply the gradients.

The code in the video provides a complete, working example. For reference, here is the core logic from the learn() function, annotated to match our theoretical discussion:

# Inside the learn() method, within a tf.GradientTape context:

# 1. Get value estimates from the Critic for current and next states
state_value, _ = self.actor_critic(state)
new_state_value, _ = self.actor_critic(new_state)

# 2. Calculate the one-step advantage (TD Error)
# Note: The video uses (1 - int(done)) to zero out the value of the next state if it's terminal.
advantage = reward + self.gamma * new_state_value * (1 - int(done)) - state_value

# 3. Calculate the Critic loss
critic_loss = advantage ** 2

# 4. Calculate the Actor loss
# Get log probability of the action that was actually taken
action_probs = tfp.distributions.Categorical(probs=probs)
log_prob = action_probs.log_prob(self.action) # self.action was saved from choose_action()

actor_loss = -log_prob * advantage # The video uses tape.stop_gradient(advantage) implicitly

# 5. Sum the losses and apply gradients
total_loss = actor_loss + critic_loss
grads = tape.gradient(total_loss, self.actor_critic.trainable_variables)
self.actor_critic.optimizer.apply_gradients(zip(grads, ...))

This implementation directly maps our A2C update rules into TensorFlow code.

4. A2C vs. A3C and Practical Enhancements

The algorithm we've built is often called A2C (Advantage Actor-Critic). It is a synchronous, deterministic version of the original A3C (Asynchronous Advantage Actor-Critic). A3C uses multiple parallel workers, each with its own copy of the environment, to gather experience asynchronously. This helps to de-correlate the agent's experience, which stabilizes learning. A2C, in contrast, often waits for a batch of experience from all workers before performing a single update. With modern GPUs, this synchronous approach is often more efficient.

To further improve stability and encourage exploration, a common addition to the A2C loss function is an entropy bonus. Entropy is a measure of randomness in the policy's action distribution. By adding the policy's entropy to the objective function, we encourage the agent to maintain a more stochastic policy and avoid collapsing to a suboptimal deterministic action too early.

The final loss function often looks like this:

where and are coefficients to weight the different loss components.

Actor Critic (A3C) Tutorial

This short video provides a concise overview of the A2C algorithm and also introduces the idea of A3C and the entropy bonus.

Watch the following clips from this video: 02:44 - 03:54: This gives a very quick summary of the A2C algorithm steps, which should reinforce what you've just learned. 03:54 - 04:58: Briefly explains A3C and the motivation for running multiple environments in parallel. 09:07 - 11:03: Focus on the loss calculation part. Note how the value_loss (Critic), policy_loss (Actor), and entropy_bonus are combined to form the final loss. This is a key practical enhancement.

Conclusion

Congratulations! You have now moved from the high-variance REINFORCE algorithm to the much more stable and powerful Actor-Critic paradigm. By training a Critic to provide a learned baseline, we can generate a much more reliable learning signal—the advantage—to guide our Actor.

Key Takeaways:

  • Actor-Critic methods combine a policy network (the Actor) and a value network (the Critic).
  • The Actor decides which actions to take, while the Critic evaluates the states visited by the Actor.
  • Advantage Actor-Critic (A2C) uses the TD error () as an estimate for the advantage.
  • The Critic is trained to minimize the TD error (e.g., ).
  • The Actor is updated using a policy gradient scaled by the advantage from the Critic (e.g., ).
  • Practical implementations often use a shared network with two heads and add an entropy bonus to the loss to encourage exploration.

Preview of the Next Lesson:

A2C is a huge step forward, but it still has a vulnerability. The policy updates are "unconstrained." A single bad batch of data could lead to a large, destructive update to the policy from which it might not recover. How can we ensure that we improve the policy without risking a catastrophic performance drop?

This is the problem solved by our next algorithm: Proximal Policy Optimization (PPO). PPO is one of the most widely used and robust RL algorithms today, and it builds directly on the Actor-Critic foundation you've established in this lesson.

Can't find a good explanation? Sign up and we'll make it for you

Sign up