Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 12
Production Systems: Concurrency and Stateful Services
1
Building a Streaming Chat API with OpenAI Compatibility
Implement an OpenAI-compatible chat completions API with streaming support on top of the custom inference engine.
2
Building an In-Memory Session Store
Design and implement an in-memory session store for managing concurrent, multi-turn conversation states.
3
Optimizing Stateful Multi-Turn Tool-Calling with KV Cache Reuse
Implement stateful multi-turn tool-calling, focusing on KV cache reuse across tool-call rounds and its impact on the scheduler.
4
API Gateway Rate Limiting with Token Bucket
Implement token-bucket rate limiting at the API gateway level.
5
Implementing a Capacity-Limited Request Queue with Graceful Degradation
Implement a request queue with a defined capacity and a graceful degradation policy (e.g., 503 Service Unavailable) for overload scenarios.
6
Optimizing API Request Latency
Profile the full end-to-end API request path to identify and optimize non-inference latency bottlenecks.
7
Deploying Containerized Inference Services
Containerize the complete inference service and deploy it.
8
Performance Testing for SLAs
Load-test the deployed service at high concurrency, measuring performance against defined SLA targets for TTFT and throughput.
End of course,
Back to course
Previous module
Capstone Project: Build a Custom Inference Engine
Next module