Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 8
High-Throughput Serving: Batching and Scheduling
1
Static Batching for LLM Inference: Limitations with Variable-Length I/O
Implement static batching for LLM inference and identify its inefficiency with variable-length inputs and outputs.
2
Designing a Continuous Batching Scheduler
Design and implement the core logic for a continuous batching scheduler that manages a queue of incoming requests.
3
Dynamic Batching for Efficient Inference
Integrate the scheduler into the inference loop, dynamically composing a batch from multiple requests at each iteration.
4
Scheduler-Cache Interface Design
Design the data structures and interface for communication between the scheduler and the paged KV cache manager.
5
Benchmarking Continuous vs. Static Batching Throughput
Benchmark the throughput (tokens/sec) of continuous batching against static batching under a mixed workload.
6
Scheduling Policies: Impact on TTFT and ITL
Implement and compare scheduling policies (e.g., FCFS, preemption) and analyze their effect on TTFT and ITL.
Previous module
Memory Optimization: Quantization
Next module
Distributed Inference for Large Models