Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI theory, architecture, models
·
Module 14
Reinforcement Learning Foundations
1
Formulating Problems as MDPs
Formulate a problem as a Markov Decision Process (MDP)
2
Solving Bellman Equations
Derive and solve the Bellman equations for value functions
3
Value and Policy Iteration
Implement value iteration and policy iteration algorithms
4
Model-Free Prediction: Monte Carlo & TD Learning
Apply model-free prediction methods: Monte Carlo and Temporal Difference (TD) learning
5
Model-Free Control: SARSA and Q-Learning
Implement model-free control algorithms: SARSA and Q-Learning
6
On-Policy vs. Off-Policy Learning
Explain the difference between on-policy and off-policy learning
Previous module
Modern Language Model Architectures
Next module
Deep Reinforcement Learning