Hello! Welcome back to our journey through modern language model architectures.
In our last lesson, we took a deep dive into Rotary Positional Embeddings (RoPE), a truly elegant solution for encoding sequence order. We saw how rotating query and key vectors, instead of adding to them, allows the attention mechanism to focus on the relative positions of tokens—a critical factor for efficiency and performance.
RoPE, however, is just one piece of the puzzle. The LLaMA family of models, which have set the standard for open-source LLMs, incorporate several other crucial architectural innovations. Today, we will analyze two of the most impactful ones: Grouped-Query Attention (GQA) and the SwiGLU activation function. Our goal is to understand how these components work and why they provide a significant edge over the original Transformer design.

Let's get started.
1. Optimizing Attention: The Memory Bandwidth Bottleneck
In our previous discussions of attention, we focused on the number of floating-point operations (FLOPs) as the primary measure of computational cost. However, during autoregressive generation (i.e., generating text one token at a time), a different bottleneck emerges: memory bandwidth.
At each generation step, the model computes attention for the new token against all previous tokens. The keys (K) and values (V) for the previous tokens are stored in a KV cache to avoid re-computation. But for every new token, the model must load the entire set of K and V heads from the GPU's high-bandwidth memory (HBM) into the much faster on-chip SRAM for processing. As the sequence grows, the size of this KV cache balloons, and the time spent just moving this data can exceed the time spent on actual computation.
LLaMA explained: KV-Cache, Rotary Positional Embedding, RMS Norm, Grouped Query Attention, SwiGLU
To understand this problem in more detail, let's watch a segment from the video "LLaMA explained" by Umar Jamil. He provides an excellent breakdown of why memory transfer becomes the limiting factor in LLM inference.
Watch from 54:02 to 57:56. Pay close attention to the comparison between a GPU's computational speed (in TFLOPs) and its memory transfer speed (in GB/s). This highlights the core issue that attention optimizations are trying to solve.
This memory bandwidth issue is what motivated a move away from the standard Multi-Head Attention (MHA) used in the original Transformer.
From Multi-Head to Grouped-Query Attention
The path to a more efficient attention mechanism involved a couple of key steps.
a) Multi-Query Attention (MQA)
The first and most direct solution proposed was Multi-Query Attention (MQA). In standard MHA, if you have query heads, you also have corresponding key heads and value heads. MQA's insight was simple: what if all query heads shared a single key and value head?
- Benefit: This dramatically reduces the size of the KV cache that needs to be loaded at each step, leading to a massive speedup in inference.
- Drawback: Forcing all query heads to use the same K and V projections is a significant reduction in the model's capacity, which can lead to a noticeable degradation in quality.
b) Grouped-Query Attention (GQA): The Best of Both Worlds
GQA, introduced in the LLaMA 2 paper, offers a pragmatic compromise. Instead of the two extremes (one K/V head for each query head in MHA, or one K/V head for all query heads in MQA), GQA creates groups of query heads. All queries within a group share a single K/V head.
For example, with 32 query heads, you could have:
- MHA: 32 K/V heads (one for each query head).
- MQA: 1 K/V head (shared by all 32 query heads).
- GQA (e.g., 4 groups): 4 K/V heads (where each K/V head is shared by a group of 8 query heads).
GQA strikes a balance: it significantly reduces the memory bandwidth requirements compared to MHA, but retains more model capacity than MQA, resulting in faster inference with minimal loss in performance.
Variants of Multi-head attention: Multi-query (MQA) and Grouped-query attention (GQA)
The video "Variants of Multi-head attention" by Machine Learning Studio has a superb visual explanation that compares MHA, MQA, and GQA side-by-side. This will help solidify your understanding of the structural differences.
Watch from 00:29 to 01:40 for a quick recap of MHA, then from 01:55 to 04:14 for the explanation of MQA, and finally from 04:14 to 07:00 to see how GQA works as a generalized version of both. The diagrams are particularly helpful.
The Implementation Perspective
From a software engineering viewpoint, GQA is implemented quite efficiently. You compute the smaller number of key and value heads and then simply "repeat" them to match the number of query heads before the attention score calculation.
The blog post "Deconstructing LLaMA 2: A PyTorch Deep Dive" shows a helper function for this.
Deconstructing LLaMA 2: A PyTorch Deep Dive
Let's look at a code snippet that makes this concrete. This blog post provides a PyTorch-centric view of the LLaMA 2 architecture.
Read the short subsection titled 'Grouped Query Attention (GQA)'. Focus on the purpose of the repeat_kv helper function. This is the core mechanism for efficiently implementing GQA in code.
The repeat_kv function essentially takes the K and V tensors, which have fewer heads, and expands them along the head dimension so they can be used in the dot-product attention with the full set of query heads. It's a memory-efficient way to get the benefits of grouping.
Test your understanding!
A large language model is designed with 40 query heads.
- In a standard Multi-Head Attention (MHA) setup, how many Key and Value heads would it have?
- In a Multi-Query Attention (MQA) setup, how many Key and Value heads would it have?
- In a Grouped-Query Attention (GQA) setup with 8 groups, how many Key and Value heads would it have? How many query heads share a single K/V pair?
Show answer
- MHA: It would have 40 Key heads and 40 Value heads, a 1:1 ratio with the query heads.
- MQA: It would have just 1 Key head and 1 Value head, shared across all 40 query heads.
- GQA: With 8 groups, it would have 8 Key heads and 8 Value heads. Each K/V pair would be shared by
40 / 8 = 5query heads.
2. A More Expressive Feed-Forward Network: SwiGLU
The second major innovation we'll discuss lies in the feed-forward network (FFN) that follows the attention block in each Transformer layer. The original Transformer used a simple two-layer MLP with a ReLU activation function in between. Modern models like LLaMA replace this with a more powerful, albeit more complex, structure using the SwiGLU activation function.

SwiGLU stands for Swish-Gated Linear Unit. Its formula is:
Let's break this down:
- The input
xis passed through two separate linear projections,WandV. - The output of the first projection,
xW, is passed through a Swish activation function. Swish itself is a smoother version of ReLU, defined as , where is the sigmoid function. - The output of the second projection,
xV, acts as a gate. - The final result is the element-wise multiplication () of the Swish-activated output and the gate output.
This gating mechanism allows the network to dynamically control the flow of information. The gate can decide, for each dimension, how much of the Swish-activated signal should pass through. This data-dependent control is more expressive than the simple on/off switch of a ReLU function.
LLaMA explained: KV-Cache, Rotary Positional Embedding, RMS Norm, Grouped Query Attention, SwiGLU
Let's return to the "LLaMA explained" video for an exploration of SwiGLU. The explanation covers the formula, performance, and a candid look at the empirical nature of deep learning research.
Watch from 01:04:05 to 01:09:39. Note the comparison to the original ReLU-based FFN and the performance improvements on various benchmarks. The humorous quote from the paper's author at the end is a great reminder that much of cutting-edge AI is driven by empirical results, not just pure theory.
While SwiGLU involves more computation than a simple ReLU (three matrix multiplications instead of two, to be fair in parameter count), the empirical evidence is clear: it consistently improves model performance, justifying the extra cost.
Conclusion
Today, we've unpacked two more of the architectural innovations that make models like LLaMA so effective. These changes represent a shift in focus from raw scale to intelligent design, optimizing the trade-offs between model performance, training cost, and, crucially, inference efficiency.
- Key Takeaways:
- Grouped-Query Attention (GQA) is a clever compromise between Multi-Head and Multi-Query Attention. It dramatically reduces the memory bandwidth required for the KV cache during inference, leading to significant speedups with minimal impact on model quality.
- SwiGLU replaces the standard ReLU in the feed-forward network with a gated activation function. This allows for more dynamic, data-dependent control over information flow, which empirically leads to better model performance.
- Both GQA and SwiGLU are prime examples of how modern LLM design is a careful balancing act, refining every component of the Transformer architecture for maximum efficiency and power.
Preview of the Next Lesson:
We've seen how to make attention and feed-forward layers more efficient. But what if we want to scale a model to trillions of parameters without having to use all of them for every single token? In our next lesson, we will explore the Mixture of Experts (MoE) architecture, a powerful technique for dramatically increasing model capacity while keeping computation constant.