Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI theory, architecture, models
·
Module 13
Modern Language Model Architectures
1
Exploring BERT and Its Variants
Analyze the architecture of BERT and its variants (RoBERTa, ALBERT)
2
GPT Architecture and Causal Attention
Analyze the architecture of GPT models and the role of causal attention
3
Scaling Laws of Large Language Models
Analyze scaling laws for large language models and their implications
4
Implementing Rotary Positional Embeddings (RoPE)
Implement advanced positional embeddings like Rotary Positional Embeddings (RoPE)
5
LLaMA's Architectural Innovations
Analyze architectural innovations in modern LLMs like LLaMA (e.g., SwiGLU, Grouped-Query Attention)
6
Scaling Model Capacity with Mixture of Experts (MoE)
Understand the Mixture of Experts (MoE) architecture for scaling model capacity
Previous module
Foundations of Language Modeling and Embeddings
Next module
Reinforcement Learning Foundations