Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own大模型后训练与智能体
Module 1
Module 2
统计学习、信息论与优化基础
Module 3
Transformer 与自回归语言模型
Module 4
监督微调的数据与训练机制
Module 5
人类反馈数据与奖励模型
Module 6
强化学习与策略梯度基础
Module 7
基于 PPO 的 RLHF 训练
Module 8
直接偏好优化与 GRPO
Module 9
后训练评测、数据飞轮与实验设计
Module 10
训练系统、显存与分布式基础
Module 11
训练框架定位与源码修改
Module 12
Agent 后训练与技术判断表达