Skip to main content
Create your own

Arithmetic Intensity of Attention and FFN Layers

Introduction

Welcome back. In our previous lesson, you learned to distinguish between compute-bound and memory-bandwidth-bound operations. We established the core decision rule: an operation's bottleneck is determined by comparing its arithmetic intensity (the ratio of computations to data movement) against the hardware's ops:byte ratio. You now have a solid conceptual understanding of why the prefill phase tends to be compute-bound and the decode phase memory-bound.

Today, we transition from the conceptual to the concrete. This lesson is dedicated to mastering the practical skill of calculating the arithmetic intensity for the key operations that form a transformer: the matrix multiplications within the attention and feed-forward network (FFN) layers. By the end of this lesson, you will be able to derive the formulas that quantify the computational density of these operations, a fundamental skill for predicting and analyzing LLM performance.

A Refresher on Calculating Intensity

Before we dive into the complexity of a transformer, let's refresh the core concept of computational (or arithmetic) intensity with a clear, concise definition and a simple example.

A very short intro to the Roofline model

The following video from NHR@FAU provides a formal definition of computational intensity and demonstrates how to calculate it for a basic vector operation.

Please watch these two short segments: What is Computational Intensity? (04:16 - 05:15): Focus on the definition of I = n / B (flops / bytes). A Concrete Example (12:48 - 14:45): Pay close attention to how the number of FLOPs and the number of bytes transferred are counted for a simple vector norm loop. This is the exact process we will apply to our transformer layers.

The key steps are always the same:

  1. Count the total floating-point operations (FLOPs). A multiply-accumulate operation, c = a * b + c, counts as 2 FLOPs.
  2. Count the total bytes of data moved between high-bandwidth memory (HBM) and the processor's on-chip memory.
  3. Divide FLOPs by bytes to get the arithmetic intensity.

Now, let's apply this methodology to the core components of a transformer.

Analyzing the Transformer's Building Blocks

A decoder-only transformer layer primarily consists of two computational blocks: a Multi-Head Attention (MHA) module and a Feed-Forward Network (FFN) module.

Multi-Head Attention and Feed-Forward Network Modules in Transformer
This diagram shows the architectural details of the Multi-Head Attention (MHA) and Feed-Forward Network (FFN) modules. Note the tensor dimensions at each step, as these are critical for calculating FLOPs and memory movement.

We will analyze the arithmetic intensity of the matrix multiplications within each of these blocks. The resource we'll use for this provides an excellent, in-depth mathematical breakdown.

All About Transformer Inference | How To Scale Your Model

The article 'All About Transformer Inference' provides the precise derivations we need. We will go through its analysis for both linear operations and the attention mechanism.

For now, just read the first three subsections under the main 'A more granular view of the Transformer' heading: A more granular view of the Transformer: This sets the stage by identifying the key operations. Linear operations: what bottlenecks us?: Follow the derivation for a generic matrix multiplication. Focus on how the arithmetic intensity simplifies to depend on the batch size B. What about attention?: Follow the derivation for the dot-product attention mechanism. Pay close attention to how the final formula for arithmetic intensity, ST/(S+T), behaves differently during prefill (S=T) versus generation (T=1).

Let's summarize the key formulas and insights from that reading.

1. FFN / Linear Layers

The FFN and the Q, K, V, O projections in attention are all fundamentally matrix-matrix multiplications. Consider a general matmul of an input activation tensor of shape [B, D] with a weight matrix of shape [D, F].

  • FLOPs: The number of multiply-accumulate operations is B * D * F. Since each is 2 FLOPs, the total is 2 * B * D * F.
  • Memory Access: We must read the input tensor (B*D), the weight matrix (D*F), and write the output tensor (B*F). Assuming bf16 (2 bytes per element), the total bytes moved are 2 * (BD + DF + BF).
  • Arithmetic Intensity (AI):

As the article points out, for typical LLM layers, the model dimensions D and F are much larger than the batch dimension B. This allows for a critical simplification: the DF term in the denominator dominates.

This is a powerful result. The arithmetic intensity of linear layers is approximately equal to the batch size of tokens being processed.

  • In the prefill phase, B is the sequence length of the prompt, which can be large (e.g., hundreds or thousands). This results in a high arithmetic intensity, making the operation compute-bound.
  • In the decode phase, we process one token at a time, so B=1. This results in a very low arithmetic intensity, making the operation memory-bound.

2. Attention Layers

The calculation for the attention mechanism is more complex but follows the same principles. The two main computations are the QK matmul and the AV matmul. Let's use the notation from the article: B is batch size, T is query sequence length, S is key/value sequence length, and D is the head dimension.

  • Total FLOPs: Approximately 4 * B * S * T * D (from 2BSTD for QK^T and 2BSTD for softmax(QK^T)V).
  • Total Memory Access: Primarily reading Q (BTD), K (BSD), V (BSD), and writing O (BTD). Total bytes are approx 2 * (2BSD + 2BTD) = 4BSD + 4BTD (using 2 bytes/element).
  • Arithmetic Intensity (AI):

This formula reveals the performance characteristics of attention:

  • Prefill (Self-Attention): The query and key sequences are the same length, so S = T.

    The arithmetic intensity scales linearly with the sequence length. For a long prompt, this becomes a high value, making prefill attention compute-bound.

  • Decode (Cross-Attention): We are generating one new token, so the query sequence has length T=1. The key/value sequence S is the length of the entire context so far.

    The arithmetic intensity is a small, constant value, regardless of how long the context grows. This is why attention during the decode phase is fundamentally memory-bandwidth-bound.

A Complete Worked Example

Let's solidify this theory by walking through a complete, end-to-end example that combines hardware specs with our newly derived formulas.

A guide to LLM inference and performance

The Baseten article 'A guide to LLM inference and performance' provides a perfect practical application of these concepts. It analyzes Llama 2 7B on an NVIDIA A10 GPU.

Please read the following sections: 'Calculating the operations per byte (ops:byte) ratio': This recaps how to find the hardware's critical ratio. 'Calculating arithmetic intensity': This introduces the plan to analyze the attention layers. 'Breaking down the attention equation': Here, they define the variables for Llama 2 7B. The two calculation blocks: Pay close attention to the total_memory_movement_in_bytes and total_compute_in_floating_point_ops calculations. They use a slightly different (but equivalent) formulation for the attention algorithm, leading to the formula 4d(N^2) + 3N^2 / 8N^2 + 8Nd. 'Discovering our inference bottleneck': This is the punchline, where the calculated arithmetic intensity is compared to the GPU's ops:byte ratio to declare the bottleneck.

As you saw in the article, they calculate an arithmetic intensity of ~62 ops/byte for Llama 2 7B's attention layer during prefill. They compare this to the A10 GPU's ops:byte ratio of 208.3. Since 62 < 208.3, they conclude the operation is memory-bound even during prefill on that specific hardware. This demonstrates that while prefill is more compute-intensive than decode, it may still be memory-bound on hardware with a very high ops:byte ratio.

Self-Assessment

To test your understanding, try to solve this problem from the JAX-ML reading.

Pop Quiz: Assume we want to take a single decode step (T=1) with a global batch size of 256 requests for a 30B parameter dense model. The model is running on a 16-chip TPU v5e slice. Each token in the KV cache is 100 KB.

Hardware specs for the entire 16-chip slice:

  • Total Memory Bandwidth: 16 * 8.1e11 Bytes/s = 12.96 TB/s
  • Total bf16 FLOPs: 16 * 1.97e14 FLOPs/s = 3.15 PFLOPs/s

The MLP blocks will be compute-bound at this batch size (B=256 is greater than the critical batch size of 240 for a TPUv5e). The attention operation will be memory-bound (since T=1).

Using the general formula for step time, calculate the expected latency:

What is a reasonable lower bound on the latency of this operation?

Click to see the solution

Let's plug in the numbers.

  • Attention part (memory-bound):

    • Batch Size = 256
    • KV Cache Size per token = 100 KB = 1e5 Bytes
    • Total Memory Bandwidth = 12.96e12 Bytes/s
    • Time = (256 * 1e5) / 12.96e12 = 1.97e-6 s ≈ 2.0 µs
  • MLP part (compute-bound):

    • Batch Size = 256
    • Parameter Count = 30e9
    • Total FLOPs/s = 3.15e15 FLOPs/s
    • Time = (2 * 256 * 30e9) / 3.15e15 = 4.87e-3 s ≈ 4.9 ms
  • Total Step Time:

    • Total Time ≈ 2.0 µs + 4.9 ms ≈ 4.9 ms

The latency is overwhelmingly dominated by the compute-bound MLP operations. The time spent on the memory-bound attention step is negligible in this high-batch scenario. Note: The JAX-ML article includes the KV cache size for the entire sequence (8192 tokens) in its calculation, which is a different scenario. Our calculation here correctly uses the per-token access cost for a single decode step.

Conclusion

In this lesson, you have moved from a qualitative to a quantitative understanding of transformer performance. You can now derive and apply the formulas for arithmetic intensity, the critical metric that governs performance bottlenecks.

Key Takeaways:

  • Arithmetic Intensity is calculated by meticulously counting FLOPs and memory accesses (FLOPs / Bytes).
  • Linear Layers (in FFN and attention projections) have an arithmetic intensity of approximately B, the token batch size. This makes them compute-bound for large batches (prefill) and memory-bound for small batches (decode).
  • Attention Mechanism has an arithmetic intensity of approximately T/2 during prefill (where T is sequence length) and ~1 during decode. This makes prefill potentially compute-bound and decode definitively memory-bound.
  • By comparing the calculated arithmetic intensity of an operation to the ops:byte ratio of the hardware, we can predict its performance bottleneck with high accuracy.

Preview of the Next Lesson:

You have now learned to calculate the arithmetic intensity (the x-axis of the Roofline plot) for key transformer operations. In the next lesson, we will put everything together. You will learn to apply the roofline model concept to determine if an operation is compute- or memory-bound on a specific GPU, visualizing where these operations fall on the plot and what that implies for optimization strategy.

Can't find a good explanation? Sign up and we'll make it for you

Sign up