Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI theory, architecture, models
·
Module 16
Aligning and Fine-Tuning Large Language Models
1
Supervised Fine-Tuning with Custom Instructions
Perform supervised fine-tuning (SFT) on a custom instruction dataset
2
LoRA for LLM Fine-Tuning
Implement LoRA for parameter-efficient fine-tuning of LLMs
3
Beyond LoRA: Exploring Other PEFT Methods
Apply other PEFT techniques like prefix and prompt tuning
4
Training a Reward Model for Human Preferences
Train a reward model to capture human preferences
5
RLHF with PPO
Implement Reinforcement Learning from Human Feedback (RLHF) using the PPO algorithm
6
Direct Preference Optimization (DPO)
Apply Direct Preference Optimization (DPO) as an alternative to RLHF
7
Evaluating Model Alignment with Human Preference Benchmarks
Evaluate model alignment using human preference benchmarks
Previous module
Deep Reinforcement Learning
Next module
LLM Interaction: Prompting and In-Context Learning