Skip to main content
Create your own

Optimizing Single-GPU Serving Performance

Introduction

In our last lesson, you gained hands-on experience with the essential tools for performance measurement, using nvidia-smi for high-level checks and torch.profiler to generate detailed memory and execution timelines. You learned how to collect the data; now, you will learn how to interpret it.

This lesson directly addresses the learning outcome: Identify performance bottlenecks in a naive single-GPU serving setup by analyzing GPU utilization and memory traces. We will move from being a data collector to a performance detective. Your goal is to analyze the traces you can now generate and pinpoint the areas where performance is being wasted.

We'll dissect the LLM inference process into its two fundamental phases—prefill and decode—and you'll learn to recognize their distinct signatures in a performance trace. Understanding these phases is the key to diagnosing why your 70B model might be slow, even with a powerful GPU.

1. The Vital Signs of LLM Inference

Before diving into a trace, we need to know what we're looking for. In LLM serving, performance is not a single number; it's characterized by a few key metrics that reflect both the user's experience and the system's efficiency.

LLM Inference Performance Engineering: Best Practices

To begin, let's establish a clear vocabulary for the key performance metrics. This blog post from Databricks provides industry-standard definitions.

Read the section 'Important Metrics for LLM Serving'. Focus on internalizing the definitions of 'Time To First Token (TTFT)' and 'Time Per Output Token (TPOT)'.

To summarize, the two most critical latency metrics are:

  • Time To First Token (TTFT): The time a user waits after sending a prompt to see the very first token of the response. A high TTFT makes an application feel sluggish and unresponsive.
  • Time Per Output Token (TPOT): The time it takes to generate each subsequent token after the first one. This is often expressed as its reciprocal, tokens per second, and determines the "speed" at which the model appears to be "typing." It's also frequently called Inter-Token Latency (ITL).

Our primary goal when analyzing a trace is to measure these two values and understand what drives them.

2. The Anatomy of a Request: Prefill and Decode

Why are TTFT and TPOT measured separately? Because they correspond to two fundamentally different computational phases of transformer inference.

Mastering LLM Inference Optimization From Theory to Cost Effective Deployment: Mark Moyou

This video by Mark Moyou provides an excellent and intuitive visualization of the entire inference workload, breaking it down into its core stages.

Watch the two segments from 02:46 to 16:36. The first part (until 14:38) explains the journey from prompt to tokens, detailing the 'prefill' and 'decode' stages and the role of the KV cache. The second part (from 14:38) explicitly defines TTFT and inter-token latency in the context of these stages.

As the video explains, every inference request involves:

  1. The Prefill Stage (or Prompt Processing):

    • The model processes all the tokens in the input prompt in parallel.
    • This involves large matrix-matrix multiplications, as the attention mechanism computes scores between all input tokens.
    • This phase is typically compute-bound, meaning its speed is limited by the raw processing power (FLOPS) of the GPU.
    • The duration of this stage is the primary driver of TTFT.
  2. The Decode Stage (or Autoregressive Generation):

    • The model generates the output one token at a time.
    • In each step, the model performs attention over all previous tokens (both from the prompt and previously generated ones) to produce the next token.
    • Thanks to the KV Cache, the model doesn't recompute attention for all previous tokens. It only needs to compute attention for the newest token against the cached keys and values.
    • This involves vector-matrix multiplications, which are typically memory-bandwidth-bound. The speed is limited not by computation, but by how fast the GPU can shuttle the model's weights from its main memory (HBM) to the on-chip compute units.
    • The duration of a single decode step determines the TPOT / ITL.

Understanding this distinction is the most critical part of building your request-to-token mental model. A slow TTFT points to a bottleneck in the compute-heavy prefill stage, while a slow TPOT points to a bottleneck in the memory-bound decode stage.

3. Practical Diagnosis: Analyzing a Performance Trace

Let's put this theory into practice. We'll use a modified version of the script from the last lesson to generate a trace and hunt for bottlenecks.

Step 1: Generate a Trace

Save the following code as diagnose_perf.py. It's similar to our previous trace_memory.py, but we'll generate more tokens to make the decode phase more prominent in our trace.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
from torch.profiler import profile, record_function, ProfilerActivity
import time

if not torch.cuda.is_available():
    raise SystemExit("No CUDA-enabled GPU found.")




# --- 1. Load Components ---
model_id = "microsoft/Phi-3-mini-4k-instruct"
print(f"Loading tokenizer and model: {model_id}...")
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    torch_dtype="auto",
    trust_remote_code=True,
)
pipe = pipeline("text-generation", model=model, tokenizer=tokenizer)
print("Model and pipeline loaded.")




# --- 2. Prepare Prompt & Generation Args ---
messages = [
    {"role": "user", "content": "Write a short story about a performance engineer debugging a slow AI model."},
]



# We'll generate a longer sequence to make the decode phase more visible
generation_args = {"max_new_tokens": 400, "return_full_text": False}




# --- 3. Warmup Run (important for stable benchmarks) ---
print("\nPerforming a warmup run...")
_ = pipe(messages, **generation_args)
print("Warmup complete.")




# --- 4. Run Inference with Profiler ---
print("\nGenerating response with profiler...")
with profile(
    activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
    profile_memory=True,
    with_stack=True,
    record_shapes=True
) as prof:
    with record_function("model_inference"):
        start_time = time.time()
        output = pipe(messages, **generation_args)
        end_time = time.time()

total_time = end_time - start_time
generated_tokens = len(tokenizer.encode(output[0]['generated_text']))
print(f"\nInference complete in {total_time:.2f} seconds.")
print(f"Generated {generated_tokens} tokens.")
print(f"Overall throughput: {generated_tokens / total_time:.2f} tokens/sec")




# --- 5. Export Trace and Print Summary ---
print("\nExporting trace...")
prof.export_chrome_trace("trace_diagnostics.json")
print("Trace exported to trace_diagnostics.json. Open it in chrome://tracing or Perfetto.")
print("\nTop 10 CUDA Kernels by Total Time:")
print(prof.key_averages().table(sort_by="self_cuda_time_total", row_limit=10))




# --- 6. Print Output ---
print("\n--- Model Output ---")
print(output[0]['generated_text'])

Run the script: python diagnose_perf.py.

This will produce a trace_diagnostics.json file. Open it in your Chrome browser by navigating to chrome://tracing and loading the file.

Step 2: Identify the Phases in the Trace

In the trace viewer, use the W/S keys to zoom in and out, and A/D to pan left and right. Find the timeline for your GPU. You should see a section of activity corresponding to the model_inference block you labeled.

You should be able to clearly distinguish the two phases:

  1. The Prefill Phase: A single, dense block of CUDA kernels at the beginning of the model_inference span. This is the model processing your input prompt. Use the measurement tool (by clicking and dragging) to find its duration. This is your TTFT.
  2. The Decode Phase: A long sequence of smaller, nearly identical, regularly-spaced groups of kernels. Each group represents one decode step (the generation of one token). Measure the time between the start of one group and the start of the next. This is your ITL/TPOT.
GPU Utilization and Memory Usage Over Time
This plot shows the GPU compute utilization (top, blue line) over time. Notice the pattern of short, intense bursts of work followed by periods of complete inactivity. This is a classic sign of a bottleneck: the GPU is spending time waiting, which indicates another part of the system (like the CPU or data pipeline) is not feeding it work fast enough.

The gaps between the decode steps in your trace are a form of this idle time. In our simple script, this "bottleneck" is the Python overhead of the Hugging Face generate loop. In a production server, it could be caused by waiting for the network, scheduling logic, or other CPU-bound tasks. The key insight is that any time the GPU is not computing, you are wasting an expensive resource.

Step 3: Analyze the Kernels

Click on the kernels within the prefill and decode blocks. While a deep dive requires NVIDIA's Nsight tools (which we'll cover in Module 4), you can already infer a lot.

  • In the prefill block, you'll likely see kernels with names related to matrix multiplication (gemm).
  • In the decode blocks, you will also see gemm kernels, but they will be operating on smaller matrices (a vector against a matrix), and you'll see more memory-related operations.

The table printed by the script (Top 10 CUDA Kernels) gives you another view, confirming which operations are consuming the most total time.

4. How Workloads Create Bottlenecks

The performance profile of your application depends heavily on the type of work it's doing. A request with a long prompt and a short required output will behave very differently from one with a short prompt that asks for a long story.

Mastering LLM Inference Optimization From Theory to Cost Effective Deployment: Mark Moyou

Let's return to Mark Moyou's presentation, where he discusses what he calls the 'most important slide'—the four querying patterns and their performance implications.

Watch the segment from 16:36 to 19:19. Pay close attention to the four quadrants he describes (long/short input vs. long/short output) and how they affect prefill time, generation time, and VRAM usage.

This framework is crucial for an AI Systems Engineer. When designing or debugging a serving system, you must ask: What is my expected workload?

  • Summarization of long documents? (Long input, short output): You'll be prefill-bound. Your main concern is a high TTFT. Optimizing the prefill stage is critical.
  • Creative writing or chatbots? (Short input, long output): You'll be decode-bound. Your main concern is a low TPOT (high tokens/sec). Optimizing the decode stage and minimizing kernel launch overhead is key.
  • RAG over a large context? (Long input, long output): The worst of both worlds. You'll need to optimize both phases and manage a large KV cache.

Your analysis of the trace_diagnostics.json file for a "short story" request should show that the total time is dominated by the long, repetitive decode phase.

Conclusion

In this lesson, you took a significant step toward thinking like a systems engineer. You've learned to move beyond just running a model and started to dissect its execution, connecting high-level performance metrics to low-level hardware behavior.

Key Takeaways:

  • LLM inference is a two-phase process: prefill (compute-bound, drives TTFT) and decode (memory-bandwidth-bound, drives TPOT).
  • Performance traces from torch.profiler clearly visualize these two phases, allowing you to measure their respective durations.
  • GPU idle time is the most obvious sign of a bottleneck. In a trace, these are "gaps" where the GPU is waiting for work from the CPU.
  • The nature of the bottleneck depends on the workload. Applications are often prefill-bound (e.g., summarization) or decode-bound (e.g., chatbots). Analyzing the expected query patterns is essential for performance tuning.

Preview of the Next Lesson:

You've now measured the VRAM used by the model and diagnosed the performance characteristics of its execution. But why does a model like Phi-3 use ~8GB of VRAM? And how would you calculate that for a 70B model? In the next module, "LLM Internals," we will start answering these questions by building a transformer from scratch. Your first task will be to write a function that calculates the total parameter count and the resulting VRAM footprint from a model's architectural hyperparameters, giving you the power to predict memory usage before ever loading a single weight.

Can't find a good explanation? Sign up and we'll make it for you

Sign up