Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 1
The Request-to-Token Mental Model
1
LLM Serving Stack: From Request to Token
Trace the full lifecycle of an HTTP request to token output by instrumenting a live LLM serving stack.
2
Inference Stages and Hugging Face `generate` Pipeline
Map the conceptual stages of inference (tokenization, prefill, decode, detokenization) to the corresponding function calls in a Hugging Face Transformers generate pipeline.
3
Profiling Inference Time with Python
Measure wall-clock time for each inference stage using Python profiling tools on a small model.
4
Profiling Prefill vs. Autoregressive Decode
Analyze profiling data to characterize the distinct compute and memory profiles of the prefill vs. autoregressive decode phases.
5
Diagnosing LLM Latency: TTFT vs. ITL
Measure and compare time-to-first-token (TTFT) vs. inter-token latency (ITL) on a running model to diagnose performance characteristics.
Previous module
Next module
Foundations: Serving and Profiling a Small Model