Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 6
Memory Optimization: Advanced KV Cache Management
1
KV Cache for Autoregressive Decoding: Memory Analysis
Implement a basic KV cache for autoregressive decoding and measure its memory growth over the generation sequence.
2
KV Cache Eviction Strategies: Impact on Generation Quality
Implement and compare different KV cache eviction strategies (e.g., sliding window, token dropping) and evaluate their impact on generation quality.
3
Prefix Caching for KV Cache Optimization
Implement prefix caching to reuse KV cache entries for requests that share a common prompt.
4
KV Cache Paging Fundamentals
Explain the memory fragmentation problem in naive KV cache allocation and the core concept of paged memory management.
5
Paged KV Cache Allocator
Implement a paged KV cache allocator that manages non-contiguous physical memory blocks for storing logical token sequences.
6
Paged KV Cache Integration & Benchmarking
Integrate the paged KV cache into an inference loop and benchmark the reduction in memory waste under a simulated concurrent workload.
Previous module
Efficient Attention Mechanisms
Next module
Memory Optimization: Quantization