Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 2
Foundations: Serving and Profiling a Small Model
1
Local PyTorch and CUDA Setup for Single GPU Inference
Set up a local inference environment with PyTorch and CUDA on a single consumer GPU.
2
Loading and Serving Small Models with Hugging Face Transformers
Load and serve a small model (e.g., Phi-3-mini) using Hugging Face Transformers.
3
Profiling PyTorch GPU Memory with nvidia-smi
Profile GPU memory usage and utilization during inference using nvidia-smi and PyTorch CUDA utilities.
4
Optimizing Single-GPU Serving Performance
Identify performance bottlenecks in a naive single-GPU serving setup by analyzing GPU utilization and memory traces.
5
Deploying and Benchmarking FastAPI Model Endpoints
Serve the model behind a FastAPI endpoint and benchmark end-to-end request latency.
Previous module
The Request-to-Token Mental Model
Next module
LLM Internals: Architecture and Memory Arithmetic