Skip to main content
Create your own

Inference Stages and Hugging Face `generate` Pipeline

Hello! Welcome back to our course.

In our last lesson, we traced the high-level journey of a request through a sophisticated serving engine like vLLM. We established a crucial mental model centered on two distinct computational phases: a compute-bound prefill phase for the initial prompt and a memory-bandwidth-bound decode phase for generating subsequent tokens, all orchestrated by a scheduler and powered by the KV cache.

Today, we will bring this high-level model down to the code level. Our goal is to map the conceptual stages of inference—tokenization, prefill, decode, and detokenization—to the corresponding function calls and logic within the Hugging Face Transformers library. While vLLM is a highly optimized system, Transformers provides the fundamental building blocks. Understanding how it works is essential before we begin building and optimizing our own systems.

By the end of this lesson, the model.generate() function will no longer be a black box. You'll understand the internal loop that executes the prefill and decode stages we discussed.

The Four Conceptual Stages of Inference

Let's begin by formally reviewing the four stages that take a text prompt and produce a text response.

  1. Tokenization: The input text string is converted into a sequence of integer token IDs that the model can process.
  2. Prefill (Prompt Processing): The model performs a single, large forward pass, processing all the input token IDs in parallel. This step is computationally intensive but highly parallelizable. Its main outputs are the logits for the first generated token and the initial state of the KV cache.
  3. Decode (Autoregressive Generation): The model enters a loop. In each iteration, it takes the token generated in the previous step and performs a forward pass to predict the next one. This process is sequential (one token at a time) and reuses the computed history stored in the KV cache to remain efficient.
  4. Detokenization: As each token ID is generated, it is converted back into human-readable text. This can happen per token (for streaming) or once at the very end.

The following video provides an excellent summary of these stages.

Understanding LLM Inference | NVIDIA Experts Deconstruct How AI Works

This video from NVIDIA experts deconstructs the inference process. Pay close attention to how they segment the workflow into distinct blocks.

Please watch these two clips: Processing Blocks (03:35 - 05:38): This section introduces the four core stages: tokenization, initial prompt processing (prefill), token generation (decode), and detokenization. Full Loop Summary (18:54 - 21:26): This clip puts everything together, showing the complete loop from tokenized prompt to detokenized output and highlighting how the prefill and decode stages fit into the overall process.

This diagram also provides a great visual summary of the end-to-end flow.

LLM Inference Flow Diagram
This diagram shows the full inference pipeline, from the user's prompt to the final output, explicitly labeling the tokenization, prefill, decode (with its iterative nature), and detokenization stages.

The "Why" Behind Prefill and Decode: Avoiding Redundancy

To understand how these stages map to code, we must first appreciate the inefficiency of a naive autoregressive loop. In such a loop, to generate token i+1, the model would re-process the entire sequence from token 0 to i. This involves massive redundant computation, as the attention calculations for tokens 0 to i-1 have already been done.

The solution is the KV Cache.

The core idea is to cache the internal attention states—the Key (K) and Value (V) vectors—for each token after they are computed.

  • During the prefill stage, we compute the K and V vectors for all prompt tokens at once and store them.
  • During each decode step, we only compute the K and V vectors for the newest token and append them to the cache. The model can then attend to the entire history efficiently by reading from this cache.

This separation of concerns—a one-time bulk computation (prefill) and a series of smaller, incremental updates (decode)—is the fundamental optimization that makes LLM inference practical.

The following article offers an exceptionally clear, code-first explanation of this process. Since you prefer practical, hands-on approaches, seeing this implemented from scratch will be highly valuable.

KV Cache from scratch in nanoVLM

The blog post 'KV Cache from scratch in nanoVLM' provides a concise, from-the-ground-up implementation of KV Caching. It perfectly illustrates the distinction between prefill and decode.

Please read the following sections. The code examples are the most important part. 'Where Redundancy Creeps In': This section uses a simple code example to prove that K and V values are needlessly recomputed in a naive loop. 'How KV Caching Fixes It': This explains the core concept of caching K and V and performing incremental updates. 'Prefill vs Decode in the Generation Loop': This is the most critical section for our lesson. It shows a clear code structure separating the PREFILL PHASE (one call with the full prompt) from the DECODE PHASE (a loop making calls with one token at a time, reusing the cache). This is the exact logic encapsulated within Hugging Face's generate method.

This diagram illustrates how the KV cache is populated during prefill and used during decode across the transformer layers.

Prefill and Decode with KV Caching in Transformer Inference
This diagram visualizes the data flow for the prefill and decode stages. During prefill, the full input sequence populates the KV cache at each layer. During decode, the query for the new token attends to the cached K and V values, which are extended with the new token's own K and V.

Mapping to the Hugging Face generate Pipeline

Now, let's connect this understanding to the transformers library. The entire process of tokenization, prefill, decoding, and detokenization is managed by a combination of the tokenizer and the model.generate() method.

Here's how the conceptual stages map to the actual function calls:

1. Tokenization
This is a straightforward call to the tokenizer object.

from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "gpt2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

prompt = "The future of AI systems engineering is"
inputs = tokenizer(prompt, return_tensors="pt")



# `inputs` is now a dictionary containing 'input_ids' and 'attention_mask'
# e.g., {'input_ids': tensor([[ 464, 2183, 286, 1754, 4016, 6499, 318]]), ...}

2. Prefill & 3. Decode
Both of these stages are handled inside the model.generate() method. The distinction between them is managed internally through an argument called past_key_values (which is often abstracted into a Cache object).

  • Prefill: The generate method begins. On its first internal forward pass, it sends the entire inputs['input_ids'] to the model. The past_key_values argument is None. The model computes the logits for the next token and, crucially, returns the populated KV cache for the entire prompt.
  • Decode: For all subsequent internal forward passes, the generate method sends only the last generated token as the new input_ids, but it also passes the past_key_values object it received from the previous step. The model uses this cache, appends the new K and V vectors, and calculates the next token.

This loop continues until a stopping condition is met (e.g., max_new_tokens is reached or an EOS token is generated).

Let's look at the transformers documentation to see where this past_key_values object lives.

Utilities for Generation

The Hugging Face 'Utilities for Generation' documentation details the inputs and outputs of the generation process. We can see the past_key_values as a key part of the model's output.

Please review these sections: 'Generate Outputs': Look at the documentation for class transformers.generation.GenerateDecoderOnlyOutput. Notice the past_key_values attribute. This is the object that carries the KV cache from one generation step to the next. 'Caches': Skim the introduction to this section and the descriptions for DynamicCache and StaticCache. This shows that past_key_values is not just a tuple but is implemented as a more sophisticated Cache object that manages the memory for the K and V tensors.

4. Detokenization
This is the final step, converting the output token IDs back to a string.




# Assuming 'outputs' is the result from model.generate()
output_ids = model.generate(**inputs, max_new_tokens=10)



# e.g., output_ids might be tensor([[ 464, 2183, ..., 198, 198, 262, 2356, ...]])




# The tokenizer's decode method handles the conversion back to text
decoded_text = tokenizer.decode(output_ids[0], skip_special_tokens=True)
print(decoded_text)



# "The future of AI systems engineering is a bright one.\n\nThe future of"

Putting It All Together

Here is a single code block that demonstrates the full mapping:

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM




# --- SETUP ---
model_id = "gpt2"
device = "cuda" if torch.cuda.is_available() else "cpu"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id).to(device)

prompt = "The future of AI systems engineering is"




# --- 1. TOKENIZATION ---
# Map text prompt to numerical token IDs
inputs = tokenizer(prompt, return_tensors="pt").to(device)

print(f"Input Token IDs: {inputs['input_ids']}")
print("-" * 20)




# --- 2. PREFILL & 3. DECODE ---
# This single call encapsulates the prefill and decode loop.
# Prefill: First forward pass with all input_ids and no past_key_values.
# Decode: Subsequent forward passes with the last token and the updated past_key_values.
output_ids = model.generate(**inputs, max_new_tokens=10, pad_token_id=tokenizer.eos_token_id)

print(f"Output Token IDs: {output_ids}")
print("-" * 20)




# --- 4. DETOKENIZATION ---
# Map the output token IDs back to a text string
decoded_text = tokenizer.decode(output_ids[0], skip_special_tokens=True)

print(f"Final Output Text: '{decoded_text}'")

Conclusion

In this lesson, we successfully bridged the gap between our high-level conceptual model of inference and the concrete implementation within the Hugging Face Transformers library. You now have a more detailed mental model of what happens inside model.generate().

Key Takeaways:

  • LLM inference consists of four stages: tokenization (tokenizer(text)), prefill, decode, and detokenization (tokenizer.decode(ids)).
  • The prefill and decode stages are both contained within the model.generate() method.
  • The key mechanism differentiating these two stages is the KV Cache, passed between internal steps via the past_key_values argument (or Cache object).
  • Prefill is the first forward pass on the full prompt, with past_key_values=None, which creates the initial cache.
  • Decode is a loop of subsequent forward passes, each using only the last generated token and the populated past_key_values from the previous step.

Preview of the Next Lesson:

We have now mapped the concepts to the code. The next logical step is to measure them. In the next lesson, we will measure the wall-clock time for each inference stage using Python profiling tools on a small model. This will be our first truly hands-on measurement exercise, turning our conceptual understanding of prefill and decode performance into hard numbers.

Can't find a good explanation? Sign up and we'll make it for you

Sign up