Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 4
Compute Optimization
1
Compute-Bound vs. Memory-Bound Transformer Inference
Distinguish between compute-bound and memory-bandwidth-bound operations in transformer inference.
2
Arithmetic Intensity of Attention and FFN Layers
Calculate the arithmetic intensity for key operations (e.g., matrix multiplications) in attention and FFN layers.
3
GPU Roofline Model for Performance Analysis
Apply the roofline model concept to determine if an operation is compute- or memory-bound on a specific GPU.
4
Optimizing QKV Projection with Operator Fusion
Implement operator fusion (e.g., for QKV projection) to reduce kernel launch overhead and improve memory access patterns.
5
Optimizing Inference with torch.compile
Apply torch.compile to the model and measure its effect on inference latency and GPU kernel execution.
6
CUDA Kernel Bottleneck Analysis with Nsight Systems
Profile inference with Nsight Systems to pinpoint compute bottlenecks at the CUDA kernel level.
Previous module
LLM Internals: Architecture and Memory Arithmetic
Next module
Efficient Attention Mechanisms