Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 5
Efficient Attention Mechanisms
1
Profiling Bottlenecks in HBM-Accessed Dot-Product Attention
Profile standard scaled dot-product attention to identify the memory-access bottleneck caused by HBM reads/writes.
2
FlashAttention: Tiling and Online Softmax
Explain the tiling and online softmax algorithm used by FlashAttention to optimize memory access.
3
FlashAttention Benchmarking
Benchmark standard attention vs. a FlashAttention implementation, measuring speedup and memory savings across varying sequence lengths.
4
Multi-Query Attention: KV Cache Reduction
Implement Multi-Query Attention (MQA) and measure the reduction in KV cache size compared to standard Multi-Head Attention.
5
From MQA to GQA: A Trade-off Analysis
Generalize the MQA implementation to Grouped-Query Attention (GQA) and analyze its trade-off between memory footprint and model quality.
Previous module
Compute Optimization
Next module
Memory Optimization: Advanced KV Cache Management