Hello! Welcome to the first lesson in our module on Deep Reinforcement Learning.
In our previous module, we established the foundations of RL, culminating in the distinction between on-policy and off-policy learning. We saw that off-policy methods like Q-learning are highly sample-efficient because they can learn from a buffer of past experiences—a technique called experience replay. However, all the methods we've used so far relied on Q-tables, which are simply not feasible for environments with large or continuous state spaces.
Today, we bridge that gap. We'll combine the power of deep neural networks with Q-learning to create an agent that can learn in much more complex environments.
This lesson directly addresses the learning outcome: Implement a Deep Q-Network (DQN) with experience replay and a target network. We will unpack the theory behind DQN and then walk through a full implementation in PyTorch.
1. From Tables to Networks: The Need for Function Approximation
The core limitation of tabular methods is the "curse of dimensionality." Imagine an environment like an Atari game where the state is represented by the screen's pixels. An 84x84 pixel grayscale image has possible states. A Q-table to store a value for each state-action pair would be astronomically large and impossible to store, let alone fill with meaningful values.
The solution is to use a function approximator. Instead of a table, we use a parameterized function to estimate the Q-values. Given your background in machine learning, you'll immediately recognize that neural networks are perfect for this role.
This leads to the central idea of a Deep Q-Network: we use a neural network that takes a state as input and outputs a vector of Q-values, one for each possible action . We can denote this as , where represents the network's weights.

Training this network becomes a regression problem: we want to adjust the weights so that the network's output gets closer to the "true" optimal Q-value. However, simply plugging a neural network into the Q-learning update rule leads to instability. The breakthrough of DQN came from two key innovations that stabilize the training process.
2. The Core Components of DQN
The original DeepMind paper, which demonstrated superhuman performance on Atari games, introduced two crucial techniques on top of the basic Q-network: Experience Replay and a Target Network.
To start, let's get a high-level conceptual overview of these components from the excellent article 'Deep Q-Networks Explained' on LessWrong.
Read 'DQN Overview (Section 3)'. Focus on understanding the purpose of the three main components: the experience replay buffer, the main neural network (acting phase), and the target network (learning phase).
Let's break down each component in more detail.
A. Experience Replay
As we briefly touched upon in the last lesson, training a neural network on highly correlated data—like consecutive frames from a game—is a recipe for disaster. The model can overfit to recent experiences and forget what it learned before.
Experience Replay solves this by creating a large buffer (often implemented as a deque in Python) that stores transitions the agent observes: . To update the network, we don't use the most recent transition. Instead, we sample a random mini-batch of transitions from this buffer.
This has two major benefits:
- Breaking Correlations: Random sampling breaks the temporal correlations in the data, making the training updates more stable and independent, similar to the i.i.d. assumption in supervised learning.
- Data Efficiency: Each experience can be reused multiple times in different training batches, allowing the agent to learn more from each interaction with the environment.
B. The Target Network
The second source of instability comes from the Q-learning update rule itself. Recall the temporal difference (TD) target from Q-learning:
We are trying to update our network's prediction, , to be closer to this target, . The problem is that the same network weights are used to compute both the prediction and the target. This is like trying to hit a target that moves every time you adjust your aim. This "chasing a moving target" problem can cause oscillations and prevent the training from converging.
The solution is to use two separate networks:
- Policy Network (): This is the main network that we are actively training. It is used to select actions (during exploitation) and to calculate the Q-value for the current state.
- Target Network (): This is a copy of the policy network. Its weights are frozen for a period of time. This network is used to calculate the value of the next state for the TD target, providing a stable, consistent target for the policy network to learn towards.
The weights of the target network () are periodically updated with the weights from the policy network (). This can be a "hard" update (a direct copy every C steps) or a "soft" update (a slow, weighted average).
This diagram provides a great visual summary of how all the components interact.

3. The DQN Algorithm and Implementation
With these concepts in place, we can now walk through a full implementation. We'll use the FrozenLake environment from the gymnasium library, which is simple enough to let us focus on the DQN mechanics.
The following video from the "Johnny Code" channel provides an exceptionally clear, step-by-step walkthrough of building a DQN in PyTorch. We will watch it in segments to understand each part of the implementation.
Simply Explaining Deep Q-Learning/Deep Q-Network (DQN) | Python Pytorch Deep Reinforcement Learning
First, let's get a conceptual walkthrough of the differences between tabular Q-learning and DQN, and how the policy and target networks are used in the training loop.
Watch from 04:03 to 14:41. This is the core conceptual part of the video. Pay close attention to: The input/output of the Q-network (04:03). The step-by-step process of using the policy and target networks to calculate the target Q-value and update the policy network (10:07). How the target Q-value is calculated differently for terminal vs. non-terminal states.
Now, let's dive into the code.
Step 1: The DQN and ReplayMemory Classes
The first two pieces we need are the neural network itself and the class for our experience replay buffer.
Simply Explaining Deep Q-Learning/Deep Q-Network (DQN) | Python Pytorch Deep Reinforcement Learning
This segment of the video shows the PyTorch implementation of the DQN neural network class and the ReplayMemory class.
Watch from 16:33 to 18:52. Notice how: The DQN class is a standard torch.nn.Module, a simple feed-forward network. The ReplayMemory class uses a Python deque with a maximum length to automatically discard old experiences. It has methods to append new experiences and sample a random batch.
For comparison, the official PyTorch DQN tutorial provides a similar implementation. Note that it uses a namedtuple for the Transition, which is a clean way to organize the stored experience.
Reinforcement Learning (DQN) Tutorial
Let's look at the official PyTorch tutorial for a slightly different, but conceptually identical, implementation of the Replay Memory.
Read the 'Replay Memory' section. Compare this implementation with the one in the video. Both achieve the same goal.
Step 2: The Training Loop
This is where everything comes together. The main training function will initialize the environment, networks, and replay buffer. Then, it will loop through episodes, and within each episode, it will:
- Select an action using an -greedy policy.
- Execute the action and observe the outcome.
- Store the transition in the replay buffer.
- Sample a mini-batch from the buffer and perform an optimization step.
- Periodically update the target network.
Simply Explaining Deep Q-Learning/Deep Q-Network (DQN) | Python Pytorch Deep Reinforcement Learning
Now for the main event: the training loop. This part of the video walks through the train and optimize functions, showing exactly how the policy network, target network, and replay buffer work together.
Watch from 18:52 to 28:26. This is the most important part of the implementation. Follow along carefully, pausing as needed. Pay special attention to: The main while loop for a single episode (the agent-environment interaction). The optimize function, which implements the core learning step. Step 4 (23:36): Passing the current states to the policy network. Step 5 (23:53): Passing the next states to the target network. Step 6 (24:11): Calculating the target Q-value using the Bellman equation. Step 7 (24:40): Updating the Q-value for the action taken. Step 8 (25:06): Calculating the loss and performing backpropagation on the policy network. Step 10 (27:10): Syncing the policy and target networks.
Test your understanding!
In the DQN algorithm, we use two networks: the policy network and the target network.
- Why do we need the target network? What would happen if we used the policy network to calculate the TD target ?
- In the
optimizefunction, we passnext_statesthrough the target network butcurrent_statesthrough the policy network. Why don't we use the target network for both?
Show answer
-
The target network is needed for stability. If we used the policy network to calculate the TD target, the target value would change at every single training step. This makes the learning process unstable, as the policy network would be "chasing a moving target." By using a target network that is updated only periodically, we provide a stable, consistent target for the policy network to converge towards.
-
The goal of the optimization step is to update the policy network. We want to adjust its weights to make its prediction for the current state more accurate. The target network's role is only to provide a stable value for the next state to compute the TD target. We don't want to update the target network during this step (its gradients are detached), and we're not trying to select actions with it—we're just using it as a stable value function approximator for .
4. A More Complex Example: Lunar Lander and Soft Updates
The FrozenLake environment has a discrete state space, which we one-hot encode. To see the true power of DQN, let's look at an environment with a continuous state space, LunarLander-v2. Here, the state is a vector of 8 continuous values (position, velocity, angle, etc.).
This also gives us a chance to introduce a common refinement: the soft update. Instead of a hard copy of weights every C steps, we slowly blend the policy network's weights into the target network at every step:
where is a small hyperparameter (e.g., 0.005). This leads to even more stable training.
AI Agent Lands Lunar on the Moon! | Deep Q-Learning | PyTorch | Reinforcement Learning | Gymnasium
This video from 'Tutorial Horizon' implements DQN for the Lunar Lander environment. We will focus on the 'learn' method, which highlights the soft update mechanism.
Watch from 21:46 to 33:34. You don't need to memorize the code, but focus on these key concepts: The overall structure is the same: sample from memory, calculate targets, calculate expected values, compute loss, and backpropagate. Bellman Equation (22:00): The video gives a clear explanation of how the target Q-value is calculated. Target Calculation (24:24): Notice how target_q_network is used on next_states, and gradients are detached (.detach()) to prevent updating the target network here. Expected Calculation (29:10): The local_q_network (policy network) is used on the states. Soft Update (31:25): Observe how the soft update is implemented at the end of the learn method, slowly blending the weights.
Conclusion
Congratulations on completing your first foray into Deep Reinforcement Learning! You've gone from the tabular world of Q-learning to a powerful deep learning-based agent.
Key Takeaways:
- DQN uses a neural network to approximate the Q-function, allowing it to handle large, high-dimensional state spaces.
- Experience Replay is crucial for stabilizing training by storing transitions in a buffer and sampling random mini-batches, which breaks temporal correlations and improves data efficiency.
- A separate Target Network provides stable TD targets, preventing the learning process from becoming unstable due to a "moving target" problem.
- The DQN loss function is typically the Mean Squared Error or Huber Loss between the target Q-value (from the target network) and the predicted Q-value (from the policy network).
- The target network can be updated with a hard copy periodically or with a soft update at each step.
Preview of the Next Lesson:
While DQN is a massive step forward, it has a known flaw: it tends to be overly optimistic and systematically overestimates Q-values. In our next lesson, we will explore why this happens and implement Double DQN, an elegant modification that helps to reduce this overestimation bias and further improve performance.