Welcome back to our capstone project. In the last lesson, we laid out the architectural blueprint for our custom inference engine, defining the four core components: the API Server, Scheduler, Cache Manager, and Model Executor. This design provides a clear separation of concerns, which is essential for building a robust and high-performance system.
Today, we transition from design to implementation. We will build the computational heart of our engine: the Model Executor. Your goal is to implement the core logic that can take a dynamically changing batch of requests—a mix of long prompts (prefill) and single-token generations (decode)—and efficiently execute a forward pass on the GPU to produce the next token for each request.
This is where the logical plan from the scheduler meets the physical reality of the hardware. Let's get started.
1. The Role of the Model Executor
Recall our architectural diagram. The Scheduler assembles a batch of requests, and the Cache Manager allocates memory for them. The Model Executor's job is to take this prepared batch, run the computation, and return the results.
The central challenge for the executor is handling the heterogeneous nature of a continuously batched workload. Unlike training, where batches are often uniform, an inference batch contains:
- Prefill requests: Sequences with many input tokens that need their KV cache populated for the first time.
- Decode requests: Sequences with just one new input token that need to generate their next token, reusing their existing KV cache.
The executor must process this mixed workload in a single, efficient forward pass.
Let's review the high-level steps involved in this process, as detailed in the vLLM team's blog post.
Inside vLLM: Anatomy of a High-Throughput LLM Inference ...
The 'Run forward pass' section of the vLLM anatomy blog post provides a precise, step-by-step breakdown of what our model executor needs to do. This will serve as our implementation guide.
Please read the section titled 'Run forward pass'. Pay close attention to the five steps listed. These steps form the core logic of our execute_model method.
As the article outlines, the executor's execute_model method can be broken down into a sequence of distinct operations: state updates, input preparation, the forward pass itself, gathering outputs, and sampling. We will implement each of these in turn.
2. Preparing Inputs for a Mixed Batch
The first task is to transform the logical batch of requests provided by the scheduler into concrete tensors that the GPU can process. A key feature of engines like vLLM is that they don't use padding. Instead, they flatten all the tokens from all requests in the batch into a single "super sequence."
Let's imagine the scheduler gives us a batch containing two requests:
- A prefill request:
[101, 872, 2304, 1198](4 tokens) - A decode request:
[345](1 token, continuing a previous generation)
The model executor must prepare the following inputs for the model's forward method:
input_ids: A 1D tensor containing all tokens concatenated together.
torch.tensor([101, 872, 2304, 1198, 345])positions: A 1D tensor specifying the logical position of each token within its original sequence.
torch.tensor([0, 1, 2, 3, 42])(assuming the decode token is at position 42 of its sequence).seq_lens: A list or tensor containing the length of each sequence before this step. This is crucial for locating the last token of each sequence after the forward pass.
[0, 42](The prefill request is new, length 0; the decode request had 42 tokens). The new lengths will be[4, 43].block_tables: A 2D tensor provided by the Cache Manager. Each row corresponds to a sequence in the batch and contains the physical block indices where that sequence's KV cache is stored. This is the "page table" that enables PagedAttention.
torch.tensor([[7, -1, -1, ...], [12, 23, 5, ...]])(Block 7 is allocated to the new prefill request; blocks 12, 23, 5 are for the existing decode request).
This flattening strategy is the essence of continuous batching at the execution level. It allows the GPU to process a single long sequence, which is highly efficient, while custom attention kernels ensure that tokens only attend to other tokens from their original sequence.

3. The execute_model Implementation
Let's start building the ModelExecutor class. We will focus on the execute_model method, which orchestrates the entire process. The implementation will be inspired by the practical code examples in the following article.
vLLM-Style Fast Inference Engine: Building from Scratch on CPU
This article provides a practical, from-scratch implementation of a continuous batching engine. We will adapt its logic for our ModelExecutor.
Review the ContinuousBatchingEngine class, particularly the process_batch and _forward_batch methods. We are essentially implementing the logic contained within these methods inside our ModelExecutor component.
Here's the skeleton of our ModelExecutor, integrating the concepts from the article:
import torch
class ModelExecutor:
def __init__(self, model, cache_manager):
self.model = model
self.cache_manager = cache_manager
# The model's lm_head is used for sampling
self.lm_head = model.get_output_embeddings()
def execute_model(self, batch: dict) -> dict[int, int]:
"""
Executes one forward pass for a batch of requests.
Args:
batch: A dictionary containing the inputs for the model, prepared by the scheduler.
Includes 'input_ids', 'positions', 'seq_lens', 'block_tables', etc.
Returns:
A dictionary mapping request_id to its newly sampled token_id.
"""
# --- Step 1: Prepare inputs (already done by scheduler, just move to GPU) ---
input_ids = batch['input_ids'].to(self.model.device)
positions = batch['positions'].to(self.model.device)
block_tables = batch['block_tables'].to(self.model.device)
# Other metadata like attention masks might be needed depending on the model
# --- Step 2: Execute the Forward Pass ---
# The model's forward method is modified to accept paged attention inputs
hidden_states = self.model(
input_ids=input_ids,
positions=positions,
block_tables=block_tables,
kv_cache=self.cache_manager.kv_cache
)
# --- Step 3: Sample the Next Tokens ---
sampled_tokens = self._sample(hidden_states, batch)
return sampled_tokens
def _sample(self, hidden_states: torch.Tensor, batch: dict) -> dict[int, int]:
# Implementation to follow
...
4. Step 2 In-Depth: The PagedAttention Forward Pass
The line hidden_states = self.model(...) is where the magic happens. This is not a standard Hugging Face forward call. For our engine to work, the model's attention layers must be replaced with a PagedAttention implementation that can:
- Accept a
block_tablefor each sequence. - Use custom CUDA kernels (like FlashAttention) to read from and write to the non-contiguous memory blocks specified in the tables.

While the implementation of the underlying kernels is complex, from the executor's perspective, it's a single function call that passes the right metadata. The hidden_states returned will be a tensor corresponding to the flattened input_ids, with a shape like [num_total_tokens, hidden_size].
5. Step 3 In-Depth: Sampling from a Flattened Output
The hidden_states tensor contains outputs for every token in the batch, but we only need to sample a next token from the last position of each sequence. This is where the seq_lens metadata becomes critical.
We can compute the indices of these last tokens. For a given set of seq_lens (the lengths of each sequence before this step) and prompt_lens (the number of new tokens we just processed for each sequence), we can find our targets.
```grasp
{
"type": "exercise",
"id": "c04f9dec-578f-4c5e-8698-d742da0428e1"
}
Conceptual logic for finding last-token indices
cumulative_prompt_lens = torch.cumsum(prompt_lens, dim=0)
last_token_indices = cumulative_prompt_lens - 1
Let's implement the `_sample` method:
```python
def _sample(self, hidden_states: torch.Tensor, batch: dict) -> dict[int, int]:
# Get the sequence data and sampling parameters from the batch
sequences = batch['sequences'] # List of (request_id, seq_len, prompt_len, sampling_params)
# 1. Calculate indices of the last token for each sequence
# The cumulative sum of prompt lengths gives the end position of each sequence's
# new tokens within the flattened hidden_states tensor.
prompt_lens = torch.tensor([s[2] for s in sequences], device=self.model.device)
last_token_indices = torch.cumsum(prompt_lens, dim=0) - 1
# 2. Gather the hidden states for these last tokens
last_token_hidden_states = hidden_states[last_token_indices]
# 3. Compute logits
logits = self.lm_head(last_token_hidden_states)
# 4. Apply sampling parameters and generate next token for each request
sampled_tokens = {}
for i, (request_id, _, _, sampling_params) in enumerate(sequences):
# For simplicity, we'll use greedy sampling (argmax) here.
# A full implementation would handle temperature, top-p, etc.
# using sampling_params.
# Get logits for the i-th sequence in the batch
sequence_logits = logits[i]
# Greedy sampling
next_token_id = torch.argmax(sequence_logits, dim=-1).item()
sampled_tokens[request_id] = next_token_id
return sampled_tokens
This _sample method correctly isolates the relevant outputs from the flattened hidden_states tensor, computes the logits, and applies a sampling strategy for each request in the batch, demonstrating how the executor handles the heterogeneous outputs.
Conclusion
In this lesson, we have implemented the functional core of our ModelExecutor. We have designed it to be the computational workhorse that understands the unique data structures of a PagedAttention-based system.
Our key takeaways are:
- The
ModelExecutor's main function,execute_model, orchestrates the entire GPU computation for a mixed batch of prefill and decode requests. - It operates on flattened input tensors (
input_ids,positions) prepared by the scheduler, which allows for highly efficient, padding-free execution. - The forward pass relies on a modified model architecture that uses PagedAttention, taking
block_tablesas input to manage non-contiguous KV cache memory. - The sampling process requires careful indexing into the flattened
hidden_statestensor to isolate the output for the last token of each sequence before applying sampling logic.
We now have the "muscle" of our engine. It can take a precisely defined workload and execute it.
Preview of the next lesson: Next, we will build the "brain." We will Integrate the continuous batching scheduler with the model executor. The scheduler will be responsible for creating the batch dictionary that our executor now knows how to process, making decisions about which requests to run based on priority and available memory.