Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 7
Memory Optimization: Quantization
1
Quantizing Models with bitsandbytes for Efficiency
Quantize a small model to INT8 using a library like bitsandbytes and compare its memory footprint and inference latency against the fp16 baseline.
2
Quantizing Models with GPTQ for INT4: VRAM and Latency Analysis
Quantize a model to INT4 using an algorithm like GPTQ and measure the resulting VRAM savings and latency change.
3
Quantization: Perplexity vs. Memory Trade-off
Compare the perplexity of INT8 and INT4 quantized models against the original to evaluate the quality vs. memory trade-off.
4
AWQ vs. GPTQ: Understanding Weight Quantization Principles
Explain the core principles of Activation-aware Weight Quantization (AWQ) and how it differs from post-training quantization methods like GPTQ.
5
Benchmarking Weight-Only vs. Weight and Activation Quantization
Benchmark a weight-only quantized model against one with both weight and activation quantization to compare their performance profiles.
6
Quantized Paged KV Cache: Memory Savings Measurement
Combine quantization with paged KV cache optimization and measure the cumulative memory savings on a single GPU.
Previous module
Memory Optimization: Advanced KV Cache Management
Next module
High-Throughput Serving: Batching and Scheduling