Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI theory, architecture, models
·
Module 15
Deep Reinforcement Learning
1
Implementing Deep Q-Networks (DQN)
Implement a Deep Q-Network (DQN) with experience replay and a target network
2
Double DQN for Overestimation Reduction
Implement Double DQN to reduce overestimation bias
3
Implementing REINFORCE
Implement the REINFORCE algorithm (a policy gradient method)
4
Actor-Critic Architectures: A2C Implementation
Build an Actor-Critic architecture (e.g., A2C)
5
PPO with Clipped Objective
Implement Proximal Policy Optimization (PPO) with clipping for stable training
Previous module
Reinforcement Learning Foundations
Next module
Aligning and Fine-Tuning Large Language Models