Introduction
In our last lesson, you built the ModelExecutor, the computational "muscle" of our inference engine. It's now capable of taking a carefully prepared, heterogeneous batch of prefill and decode tasks and executing them in a single, efficient forward pass.
Today, we build the "brain" that directs that muscle. This lesson focuses on fulfilling a core objective of this module: integrating a continuous batching scheduler with the model executor. You will design and implement the Scheduler component, the central orchestrator responsible for maximizing GPU utilization and system throughput. The scheduler's primary role is to observe incoming requests, manage their lifecycle, and dynamically decide which requests to run at every single iteration, creating the batches that our executor now knows how to process.
By the end of this lesson, you will have connected the main components of our engine, transforming them from isolated parts into a cohesive system that can handle a continuous stream of concurrent requests.
1. The Power of Iteration-Level Scheduling
To appreciate the scheduler's role, we must first understand the problem it solves. Traditional "static" batching is highly inefficient for LLMs. It groups requests, runs them to completion, and only then starts the next batch. If one request is long, all others wait, leaving the GPU idle.
Continuous batching, also known as iteration-level scheduling, solves this. The core idea is simple but powerful: the batch is re-formed at every single step of the model's execution. This allows the system to constantly pack the GPU with work, mixing and matching parts of different requests.
To see why this is so critical for performance, let's watch a segment from the OSDI presentation on Orca, a pioneering serving system that introduced this concept.
OSDI '22 - Orca: A Distributed Serving System for Transformer-Based Generative Models
The researchers behind Orca clearly articulate the limitations of older, request-level batching systems and introduce iteration-level scheduling as the solution. This will ground our understanding of the 'why' behind continuous batching.
Watch the sections explaining the problems with current systems and Orca's solution of iteration-level scheduling. Pay attention to how dynamic batching prevents short requests from getting stuck behind long ones and allows new requests to be added immediately.
This principle of dynamically composing a batch at each iteration is the foundational concept for our scheduler. It enables the engine to achieve very high GPU utilization, which directly translates to higher throughput.
2. The Scheduler's Architecture and Responsibilities
The scheduler sits at the heart of our inference engine, acting as the central coordinator. It makes the critical decisions that drive the entire system forward.
Let's look at a diagram that places the scheduler within the vLLM architecture.

As the diagram and the resources from previous lessons suggest, the scheduler has several key responsibilities. The nano-vLLM guide provides an excellent summary of these.
nano‑vLLM: Build‑from‑Scratch Architectural Labs
The 'nano-vLLM' guide clearly breaks down the scheduler's role. Let's review its primary responsibilities to structure our implementation.
Please read Section '3.2.2 Core Technical Analysis' and Section '6.1.1 Core Components & Responsibilities', focusing on the description of the Scheduler. Note its four primary duties: Request State Management, Batching Decision, Resource Allocation, and Policy Enforcement.
To summarize, our scheduler will:
- Manage Request State: Track requests as they move from a
waitingqueue (newly arrived) to arunningqueue (actively being processed), and finally to acompletedstate. - Make Batching Decisions: In each iteration, select which requests from the
waitingandrunningqueues will be included in the next forward pass. This decision is constrained by a token budget to keep iteration latency predictable. - Allocate Resources: Communicate with the
CacheManagerto ensure KV cache blocks are available and allocated for the selected requests. - Enforce Policy: Decide the order of execution. For now, we'll use a simple First-Come-First-Served (FCFS) policy, with a crucial twist: decode requests (running) are prioritized over prefill requests (waiting). This minimizes latency for users already receiving tokens.
3. The Scheduling Loop
The core of the scheduler is its main loop, which we can implement as a step() method. This method is called repeatedly and performs three main phases: schedule, execute, and post-process.
The "Inside vLLM" blog post provides the most detailed breakdown of this process.
Inside vLLM: Anatomy of a High-Throughput LLM Inference System
Let's dive into the specifics of the scheduling logic. The 'Inside vLLM' blog post details the engine's step-by-step execution loop, which we will replicate.
Please read the sections titled 'Generate function' and 'Scheduler'. Focus on the three stages of the step() function (Schedule, Forward pass, Postprocess) and the logic for prioritizing decode requests from the running queue over prefill requests from the waiting queue.
Let's translate this logic into the Python implementation for our Scheduler class.
Implementation: The Scheduler Class
We'll define a Scheduler class that holds the request queues and orchestrates the step cycle. It will be initialized with the other major components it needs to command: the ModelExecutor and the CacheManager.
from collections import deque
class Scheduler:
def __init__(self, executor, cache_manager, max_token_budget):
self.executor = executor
self.cache_manager = cache_manager
self.max_token_budget = max_token_budget
# Queues for managing request lifecycle
self.waiting_queue = deque()
self.running_queue = []
self.completed_requests = {}
def add_request(self, request):
"""Adds a new request to the waiting queue."""
self.waiting_queue.append(request)
def step(self):
"""Performs one iteration of the scheduling, execution, and post-processing cycle."""
# 1. Schedule a batch of requests to run
scheduled_requests = self._schedule()
if not scheduled_requests:
return # Nothing to process
# 2. Prepare the batch for the model executor
batch = self._prepare_batch(scheduled_requests)
# 3. Execute the model forward pass
sampled_tokens = self.executor.execute_model(batch)
# 4. Post-process the results
self._postprocess(scheduled_requests, sampled_tokens)
def _schedule(self):
# Implementation to follow
...
def _prepare_batch(self, scheduled_requests):
# Implementation to follow
...
def _postprocess(self, scheduled_requests, sampled_tokens):
# Implementation to follow
...
4. Implementing the Scheduling Logic
The _schedule method is where the core decision-making happens. It constructs a list of requests for the current iteration based on priority and the available token budget.

Here is the implementation of the _schedule method, following the logic from our resources.
def _schedule(self):
"""
Selects requests for the next batch, prioritizing running over waiting.
Also handles moving requests from waiting to running.
"""
current_token_budget = self.max_token_budget
scheduled_requests = []
# --- Stage 1: Prioritize running (decode) requests ---
# Sort running requests (e.g., by arrival time) if needed, but for now we iterate.
# We use a temporary list to rebuild the running_queue.
next_running_queue = []
for req in self.running_queue:
if current_token_budget >= 1: # Each decode step consumes 1 token
# Check if Cache Manager can allocate one more slot
if self.cache_manager.can_allocate(req, 1):
self.cache_manager.allocate(req, 1)
scheduled_requests.append(req)
current_token_budget -= 1
next_running_queue.append(req) # Keep in running queue
else:
# Not enough memory, preempt or pause (for now, just keep waiting)
next_running_queue.append(req)
else:
next_running_queue.append(req) # No budget, will run in next iteration
self.running_queue = next_running_queue
```grasp
{
"type": "exercise",
"id": "bb82053d-d007-4077-882c-751c06b73a47"
}
# --- Stage 2: Schedule new (prefill) requests from the waiting queue ---
while self.waiting_queue and current_token_budget > 0:
req = self.waiting_queue[0] # Peek at the first waiting request
prompt_len = len(req.prompt_token_ids)
if prompt_len <= current_token_budget:
# Check if Cache Manager has enough free blocks for the whole prompt
if self.cache_manager.can_allocate(req, prompt_len):
# Move from waiting to running
req = self.waiting_queue.popleft()
self.running_queue.append(req)
self.cache_manager.allocate(req, prompt_len)
scheduled_requests.append(req)
current_token_budget -= prompt_len
else:
# Not enough memory for this new request, stop scheduling prefill
break
else:
# This request is too long for the current budget, stop scheduling prefill
# A more advanced scheduler could implement chunked prefill here
break
return scheduled_requests
### 5. Integrating the Full Cycle
With the scheduling decision made, the scheduler's `step` method continues by preparing the batch for the executor and processing the results.
1. **`_prepare_batch`**: This method takes the `scheduled_requests` and constructs the dictionary that the `ModelExecutor` expects. It gathers `input_ids`, `positions`, `block_tables` (from the `CacheManager`), and other metadata, flattening them into the single "super sequence" format we defined in the last lesson.
2. **`_postprocess`**: After `execute_model` returns the `sampled_tokens`, this method updates the state of each request.
* It appends the new token to the request's generated sequence.
* It checks for stop conditions (e.g., EOS token, `max_tokens` reached).
* If a request is finished, it moves it to `completed_requests` and, crucially, tells the `CacheManager` to **free the KV cache blocks** associated with that request. This recycling of memory is what allows new requests to be scheduled.
The integration is now complete at a logical level. The `Scheduler` acts as the system's main loop, taking incoming requests, making intelligent decisions about how to batch them, commanding the `Executor` to perform the computation, and managing the `CacheManager` to orchestrate memory usage.
### Conclusion
In this lesson, you have designed and implemented the central orchestrator of our inference engine. You've brought together the components we've built, enabling a true continuous batching system.
**Key Takeaways:**
* The **Scheduler** is the "brain" of the inference engine, implementing **iteration-level scheduling** to maximize GPU utilization.
* It manages request lifecycles using **`waiting` and `running` queues**, prioritizing active (decode) requests to ensure low inter-token latency.
* The core **`step()`** loop consists of three phases: **scheduling** (deciding what to run based on policy and budget), **execution** (passing the batch to the executor), and **post-processing** (updating states and freeing resources).
* The integration between the Scheduler, Executor, and Cache Manager forms a feedback loop: the scheduler allocates resources, the executor uses them, and upon completion, the scheduler frees them for the next cycle.
**Preview of the next lesson:**
We have established that the scheduler must communicate with the cache manager (e.g., `can_allocate`, `allocate`, `free`). In the next lesson, we will formalize this interaction by designing the specific data structures and a robust API for the `PagedKVCacheManager`. This will solidify the memory management backbone of our engine, ensuring it is both efficient and reliable.