Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 3
LLM Internals: Architecture and Memory Arithmetic
1
Building a Decoder-Only Transformer in PyTorch
Implement a minimal decoder-only transformer from scratch in PyTorch, including embedding, attention, FFN, and output layers.
2
Calculating Transformer Parameters
Write a function to programmatically compute the total parameter count of a transformer from its architectural hyperparameters.
3
VRAM for Model Weights: Parameter Count & Data Type
Calculate the VRAM required to store model weights for a given parameter count and data type (fp32, fp16, bf16).
4
Activation Memory Calculation
Calculate the activation memory consumed during a forward pass for a given batch size and sequence length.
5
Decoder Block: Compute & Memory Analysis
Analyze the forward pass of a decoder block, mapping operations to their respective compute (MatMul) and memory (data movement) costs.
6
KV Cache Size Calculation
Calculate the total KV cache size for a given model architecture, batch size, and sequence length.
7
Demystifying 70B Model Memory: VRAM & KV Cache
Apply VRAM and KV cache calculations to a 70B model to explain its large memory footprint.
Previous module
Foundations: Serving and Profiling a Small Model
Next module
Compute Optimization