Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 11
Capstone Project: Build a Custom Inference Engine
1
Component Architecture and Interface Design
Design the high-level component architecture (API server, scheduler, executor, cache manager) and define their interfaces.
2
Core Model Executor: Prefill and Decode in Continuous Batches
Implement the core model executor, capable of running prefill and decode steps for a continuously changing batch of requests.
3
Integrating Continuous Batching in Model Execution
Integrate the continuous batching scheduler with the model executor.
4
Integrating Paged KV Cache Manager
Integrate the paged KV cache manager into the engine's memory management and execution loop.
5
Optimized Model Integration with Quantization
Integrate quantized model loading and execution into the engine.
6
Tensor-Parallel Execution for Multi-GPU Deployment
Add tensor-parallel execution capabilities to the engine for multi-GPU deployment.
7
Building Request & Response Layers for E2E Tests
Implement request ingestion and response streaming layers for an end-to-end test.
8
Benchmarking Custom Engine vs. vLLM: Performance and Architectural Analysis
Benchmark the custom engine against vLLM under a realistic traffic pattern and analyze its architectural strengths and weaknesses.
Previous module
Evaluating Production Inference Frameworks
Next module
Production Systems: Concurrency and Stateful Services