Introduction
Welcome to Module 4: Compute Optimization. In the previous module, we developed a precise "napkin math" for the memory footprint of large language models. You learned how to account for every gigabyte of VRAM consumed by model weights, the KV cache, and activations. This answered the question of why LLMs are so memory-intensive.
Now, we shift our focus from memory capacity to time and speed. This module tackles the other side of the performance coin: why LLM inference can be slow, and what we can do about it. The foundational concept for all compute optimization is understanding the primary bottleneck at any given moment.
Today's lesson addresses this directly. Our learning outcome is to distinguish between compute-bound and memory-bandwidth-bound operations in transformer inference. Mastering this distinction is the first and most critical step in diagnosing performance issues and applying the correct optimization strategy. You will learn to move beyond treating the GPU as a black box and start reasoning about its fundamental architectural limits.
The Two Phases of Inference and Their Bottlenecks
As you know from our previous discussions, LLM inference isn't a single, monolithic process. It's composed of two distinct phases: prefill and decode. Intuitively, you might guess they have different performance characteristics, and you'd be right.
Let's start with a short, clear explanation of these two phases and why their computational profiles differ.
AI Optimization Lecture 01 - Prefill vs Decode - Mastering LLM Techniques from NVIDIA
This video from Faradawn Yang provides an excellent high-level overview of the prefill and decode stages, highlighting why one is compute-heavy and the other is memory-heavy.
Please watch from 01:18 to 03:16. Focus on how the video characterizes the computation in each phase: prefill involves large, parallel matrix multiplications, while decode involves small computations but relies on accessing a large, growing memory state (the KV cache).
The key insight is that the nature of the work being done changes dramatically between processing the prompt and generating subsequent tokens.

To formalize this, we need to define our terms. An operation is:
- Compute-bound when the total time is dominated by arithmetic calculations. The GPU's processing units (the CUDA cores and Tensor Cores) are the limiting factor, and they are running at or near 100% utilization. The speed is limited by the number of Floating Point Operations Per Second (FLOPS) the GPU can execute.
- Memory-bandwidth-bound when the total time is dominated by moving data (e.g., weights, activations, KV cache) between the GPU's main memory (HBM) and its on-chip processing units (SRAM/registers). The processing units are often idle, waiting for data to arrive. The speed is limited by the memory bandwidth (in GB/s) of the GPU.
To build a robust mental model, it's essential to understand the underlying hardware.
GPU Architecture and Operation Types
The following article, 'Transformers Inference Optimization Toolset' from AstraBlog, provides an excellent overview of GPU architecture and formally defines compute- and memory-bound operations. This context is crucial for the rest of the module.
Please read the first two sections: 'GPU architecture overview': Focus on the memory hierarchy (HBM, L2, L1/SRAM) and the relative cost of accessing each level. 'Arithmetic intensity vs ops:byte': This section provides the formal definitions of compute-bound and memory-bound operations. Pay close attention to the introduction of arithmetic intensity.
Quantifying the Bottleneck: The Roofline Model
The article you just read introduced the core concept for quantifying performance bottlenecks: arithmetic intensity. It is the ratio of compute operations to memory accesses for a given algorithm.
Arithmetic Intensity = Total FLOPs / Total Bytes Accessed
We can determine the bottleneck by comparing the arithmetic intensity of our algorithm to a similar ratio for our hardware.
This hardware ratio, often called the ops:byte ratio, tells us how many operations the GPU can perform in the time it takes to move one byte of data from HBM.
ops:byte Ratio = Peak FLOPS / Memory Bandwidth (GB/s)
Let's see a practical example of how to find this hardware ratio.
Calculating the Operations per Byte Ratio
The Baseten article, 'A guide to LLM inference and performance,' walks through this calculation for a specific GPU, the NVIDIA A10.
Read the sections 'Reading GPU specs' and 'Calculating the operations per byte (ops:byte) ratio'. This will show you exactly which numbers to pull from a GPU datasheet and how to combine them to get the hardware's critical ops:byte ratio.
Now we have our decision rule:
- If
Arithmetic Intensity < ops:byte Ratio, the operation is Memory-Bound. The algorithm performs too few computations for each byte it fetches, so the GPU spends most of its time waiting for data. - If
Arithmetic Intensity > ops:byte Ratio, the operation is Compute-Bound. The algorithm performs many computations for each byte, fully utilizing the GPU's arithmetic units.
This relationship is visually captured by the Roofline Model, a fundamental tool in performance engineering.

Analysis of Transformer Inference Phases
Let's apply this quantitative framework to the prefill and decode phases.
1. Prefill Phase (Compute-Bound)
During prefill, the model processes the entire input sequence of length in parallel. The self-attention mechanism, a dominant part of the computation, involves large matrix multiplications like , where both matrices have dimensions related to .
The number of compute operations scales roughly with , while the memory accesses scale with .
Arithmetic Intensity of Multi-Head Attention
The AstraBlog article you read earlier contains a precise analysis of the arithmetic intensity for Multi-Head Attention, which is characteristic of the prefill phase.
Read the subsection on Multi-head Attention. The key takeaway is the final formula, which shows that as sequence length L grows, the arithmetic intensity increases, often exceeding the ops:byte ratio of the hardware. This pushes the operation into the compute-bound regime.
For long input sequences, the sheer volume of parallelizable computation (large matrix-matrix products) outweighs the data movement costs, saturating the GPU's compute units.
2. Decode Phase (Memory-Bandwidth-Bound)
During autoregressive decoding, the situation is reversed. At each step, we generate a single token. The computation involves processing this one token's embeddings against the entire KV cache from all previous tokens.
- Compute: The amount of new computation is small, scaling with . We are essentially performing matrix-vector multiplications.
- Memory Access: The amount of data moved is enormous. We must read the full and tensors (the KV cache) from slow HBM into the fast on-chip SRAM for every single token we generate. The size of the KV cache is , where is the number of layers.
This combination of low compute and high memory access results in a very low arithmetic intensity.
Arithmetic Intensity of Decoding with KV Cache
Let's return to the AstraBlog article one last time for its analysis of the decoding phase.
Now, read the subsection on the KV Cache. Notice how the arithmetic intensity calculation changes dramatically. The conclusion that the intensity is less than 1, and therefore definitely smaller than the hardware's ops:byte ratio, is the critical insight. This firmly places the decode phase in the memory-bound regime.
This fundamental difference is also highlighted in research on disaggregated serving systems.
DistServe: disaggregating prefill and decoding for goodput-optimized LLM inference
This clip from a talk on DistServe at the PyTorch Conference succinctly confirms our findings from a systems perspective.
Watch from 05:11 to 06:54. The speaker explicitly states that prefill is compute-bound and can saturate a GPU even with a small batch, while decoding is memory-bound and needs a much larger batch size to approach the compute limit.
Conclusion
In this lesson, you've learned to analyze LLM inference performance through the lens of its fundamental hardware limitations. This is a significant step up from simply measuring wall-clock time.
Key Takeaways:
- GPU operations are limited by one of two factors: compute performance (FLOPS) or memory bandwidth (GB/s).
- The Roofline Model provides the theoretical framework for determining the bottleneck by comparing an algorithm's arithmetic intensity to the hardware's
ops:byteratio. - Transformer inference has two distinct phases with different bottlenecks:
- Prefill involves large, parallel matrix multiplications, making it compute-bound.
- Autoregressive Decoding involves small computations but requires reading the large KV cache from memory at every step, making it memory-bandwidth-bound.
This distinction is not just academic; it dictates our entire optimization strategy. Trying to optimize a memory-bound operation by making the math faster is pointless if the GPU is just waiting for data. Conversely, optimizing data movement for a compute-bound operation will yield no speedup.
Preview of the Next Lesson:
Now that you can identify the bottleneck, we will dive deeper into practical analysis. In the next lesson, you will learn to calculate the arithmetic intensity for key operations within the attention and FFN layers. We will then apply the roofline model concept to determine if a specific operation is compute- or memory-bound on a specific GPU, moving from the general (prefill/decode) to the specific (individual matrix multiplications, element-wise operations).