Welcome to the capstone project! Over the next set of lessons, you will synthesize all the concepts we've covered—from memory calculations and kernel optimizations to advanced batching and parallelism—to build your own high-performance LLM inference engine from the ground up.
This first lesson is all about architecture. Before writing a single line of implementation code, a systems engineer must first create a solid blueprint. We will design the high-level component architecture for our engine, defining the fundamental building blocks and how they communicate. By the end of this lesson, you will have a clear mental model and a formal design for the system you are about to build, specifically by identifying the roles of the API server, scheduler, model executor, and KV cache manager, and defining their interfaces.
1. The Anatomy of a Modern Inference Engine
A high-performance inference engine is a complex distributed system, even when running on a single GPU. To manage this complexity, we'll apply the principle of separation of concerns, breaking the system into logical, modular components. While different frameworks have unique implementations, most modern engines are built around a few core ideas.
Let's start by looking at the high-level architecture of vLLM, a state-of-the-art inference engine.

Based on this and other production systems, we can identify four primary components for our engine:
- API Server: The public-facing entry point that handles user interaction.
- Scheduler: The control unit that decides which requests to process and when.
- Model Executor: The computational workhorse that runs the model's forward pass.
- Cache Manager: The specialized memory allocator for the KV cache.
Let's explore the responsibilities and design of each of these components.
2. Component Deep Dive: Responsibilities and Interfaces
A robust architecture is defined not just by its components, but by the clarity of their interfaces. As we explore each component's role, think about the inputs it requires and the outputs it produces.
a. The API Server
The API Server is the gateway to our engine. It abstracts the entire complexity of the inference process behind a simple, familiar interface, typically a REST API compatible with OpenAI's standards.
-
Responsibilities:
- Expose HTTP endpoints (e.g.,
/v1/chat/completions). - Receive and validate incoming requests (prompts, sampling parameters, etc.).
- Tokenize raw text prompts into token IDs.
- Assign a unique ID to each request and package it into a standardized internal format.
- Forward the request to the core engine (specifically, the scheduler).
- Receive generated tokens from the engine, detokenize them, and stream them back to the client.
- Expose HTTP endpoints (e.g.,
-
Architectural Considerations:
The API server's primary function is I/O-bound (network communication, tokenization). The core engine, however, is compute-bound. A critical design decision is to decouple the API server from the model executor. They should run in separate processes or, at a minimum, separate threads, communicating asynchronously via queues. This prevents slow clients or network latency from blocking the GPU, which is the most expensive resource.

b. The Scheduler
The Scheduler is the "brain" of the inference engine. Its main job is to maximize throughput and GPU utilization by intelligently deciding which requests to run in each iteration of the main engine loop.
To understand the scheduler's role, let's watch a short segment from a video that visually breaks down the inference process in vLLM.
How the VLLM inference engine works?
This video provides an excellent visual mental model for the life of a prompt inside an inference engine. Pay close attention to the distinction between the 'prefill' and 'decode' stages, as this is the fundamental concept the scheduler manages.
Watch from 00:21:54 to 00:29:00. Focus on how the scheduler deals with two types of tasks: processing the initial prompt (prefill) and generating subsequent tokens (decode). Notice the mention of prioritizing decoding requests.
As the video explained, the scheduler manages two distinct workloads:
-
Prefill: A compute-intensive pass over the user's entire prompt to generate the very first token and populate the KV cache.
-
Decode: A memory-bandwidth-intensive pass using a single new token to generate the next one, heavily relying on the already-computed KV cache.
-
Responsibilities:
- Maintain queues for incoming requests (
waiting) and requests being actively processed (running). - Implement a policy for selecting requests for the next batch. A common policy is to prioritize
decoderequests to keep ongoing generations flowing, and then fill the remaining capacity withprefillrequests. This is the core of continuous batching. - Interact with the Cache Manager to check for available memory and allocate KV cache blocks for the selected requests.
- Assemble the final batch of requests to be sent to the Model Executor.
- Maintain queues for incoming requests (
c. The Cache Manager
The Cache Manager is a highly specialized memory management unit responsible for one thing: the KV cache. In modern engines using PagedAttention, this component is the heart of memory efficiency.
To understand its mechanics, let's turn back to the same video, which now explains how memory is organized into "blocks".
How the VLLM inference engine works?
This next segment details how the KV cache is physically stored in VRAM. Focus on the concept of 'blocks' as the unit of memory allocation and the 'free block queue' that the cache manager uses to track available memory.
Watch from 00:29:00 to 00:40:48. This will give you a concrete understanding of how the GPU VRAM is partitioned and what a KV cache block represents. You will see how prompts are mapped to these blocks and how the total VRAM budget dictates the number of available blocks.
- Responsibilities:
- At startup, partition the available GPU VRAM into a large number of fixed-size physical memory blocks.
- Maintain a pool of available blocks (the
free_block_queue). - Provide functions to
allocateandfreeblocks for a given request. - Manage the logical-to-physical mapping for each request. This mapping, often called a block table or page table, tracks which physical blocks hold the KV cache for a specific sequence of tokens.
The Cache Manager abstracts away the details of memory fragmentation, allowing the Scheduler to simply ask, "Can I have N blocks for this request?" without worrying about where they are in VRAM.
d. The Model Executor
The Model Executor is the "muscle." It takes a batch prepared by the Scheduler and orchestrates the actual computation on the GPU.
- Responsibilities:
- Prepare the inputs for the forward pass on the GPU (e.g.,
input_ids,positions,block_tables). - Execute the model's forward pass, which involves launching a sequence of CUDA kernels. This includes the paged-attention kernels that use the
block_tablesto read/write to the correct KV cache locations. - After the forward pass, perform sampling on the resulting logits to generate the next token for each request in the batch.
- In a distributed setting (which we'll implement later), the executor also manages the communication between GPUs (e.g.,
all-reducefor Tensor Parallelism).
- Prepare the inputs for the forward pass on the GPU (e.g.,
To see how these components fit together in a complete system, let's study the vLLM architecture in more detail.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System
This blog post, 'Inside vLLM: Anatomy of a High-Throughput LLM Inference System', provides a definitive guide to the architecture we are modeling. It formalizes the concepts we've just discussed.
Read the following sections: 'LLM Engine & Engine Core', 'Generate function', 'Scheduler', and 'Run forward pass'. As you read, map the text to the four components we've defined (API Server, Scheduler, Cache Manager, Executor). Notice how the LLM Engine class acts as a container for these components and how the step() function orchestrates their interaction.
3. Defining the Component Interfaces
Now, let's formalize the design by defining the interfaces between these components. We'll use Python-like class definitions to sketch out the key methods and data structures. This "header file" approach is a common practice in systems design.
# --- Data Structures ---
class SamplingParams:
# temperature, top_p, top_k, max_tokens, stop_strings, etc.
...
class Request:
request_id: str
prompt: str
prompt_token_ids: list[int]
sampling_params: SamplingParams
arrival_time: float
# Internal state: status (WAITING, RUNNING, FINISHED), generated_tokens, etc.
...
class Batch:
# A collection of data needed by the executor for one forward pass.
# Includes flattened token_ids, positions, block_tables, etc. for all requests.
...
# --- Components ---
class CacheManager:
def __init__(self, total_gpu_blocks: int, block_size: int): ...
def allocate_blocks(self, num_blocks: int) -> list[int]:
"""Allocate a specified number of physical blocks from the free pool."""
...
def free_blocks(self, block_indices: list[int]):
"""Return physical blocks to the free pool."""
...
def get_num_free_blocks(self) -> int:
...
class ModelExecutor:
def __init__(self, model_config, parallel_config): ...
def execute_model(self, batch: Batch) -> dict[str, int]: # {request_id: next_token_id}
"""Runs a single forward pass and returns the next token for each request."""
...
class Scheduler:
def __init__(self, config, cache_manager: CacheManager, executor: ModelExecutor):
self.waiting_queue: list[Request] = []
self.running_queue: list[Request] = []
self.cache_manager = cache_manager
self.executor = executor
def add_request(self, request: Request):
"""Add a new request to the waiting queue."""
...
def schedule(self) -> dict[str, int]: # {request_id: next_token_id}
"""
1. Create a batch of requests from waiting/running queues based on policy.
2. Allocate KV cache blocks via the CacheManager.
3. Prepare the Batch object.
4. Pass the batch to the ModelExecutor.
5. Return the sampled tokens.
"""
...
class APIServer:
def __init__(self, scheduler: Scheduler):
self.scheduler = scheduler
async def handle_completion_request(self, http_request):
"""
1. Validate request and create a Request object.
2. Add the request to the scheduler via self.scheduler.add_request().
3. (In the background) the engine's main loop will call scheduler.schedule().
4. Receive generated tokens and stream them back to the client.
"""
...
```grasp
{
"type": "exercise",
"id": "f78f9281-7a89-4698-a0ee-f41f1a59c942"
}
--- Main Engine Loop (Conceptual) ---
This loop would run in its own process/thread.
scheduler = ...
while True:
# The schedule() method contains the core logic of a single engine step.
sampled_tokens = scheduler.schedule()
# Process outputs, update request states, etc.
...
This blueprint clearly defines the separation of concerns and the flow of data. The `APIServer` talks to the `Scheduler`. The `Scheduler` orchestrates the `CacheManager` and `ModelExecutor`. The main engine loop simply drives the `Scheduler` forward, step by step.
### 4. A Comparative Perspective
This architectural pattern is common, but the implementation details vary. Different design choices lead to different performance trade-offs.
* **TensorRT-LLM**: NVIDIA's framework emphasizes this modularity, providing C++ implementations of a `KVCacheManager` and `Scheduler` that can be driven by a Python runtime. This highlights the performance-critical nature of these components.
* **SGLang**: This engine introduces a `Router` component that sits in front of the schedulers. This router is "cache-aware"—it tries to send requests to a specific worker node that might already have the request's prefix cached, further optimizing performance in a distributed setting. Its `RadixAttention` replaces the block-based `CacheManager` with a more complex tree structure for automatic prefix sharing.
This shows that while our four components are a solid foundation, there's always room for innovation in their design and interaction.
```grasp
{
"type": "exercise",
"id": "94a64e04-347e-4b33-866b-e21fade8a500"
}
Conclusion
In this lesson, we designed the high-level architecture for our custom inference engine. We established that a modular design is essential for managing complexity and enabling performance.
Our key takeaways are:
- An inference engine can be broken down into four core components: API Server, Scheduler, Model Executor, and Cache Manager.
- Each component has a distinct responsibility, from handling network I/O to orchestrating GPU computation.
- The interfaces between these components must be clearly defined. The
Scheduleracts as the central orchestrator, coordinating the other components within the main engine loop. - This architecture is a common pattern, but frameworks like SGLang and TensorRT-LLM innovate on the internal design of these components to achieve different performance goals.
With this architectural blueprint in hand, we are now ready to start building.
Preview of the next lesson: In the next lesson, we will begin implementing our design. We'll start with the computational core: Implement the core model executor, capable of running prefill and decode steps for a continuously changing batch of requests.