Hello! Welcome to the next lesson in our journey through the Transformer architecture.
In the previous lesson, we constructed the Transformer's encoder block. We saw how it uses multi-head self-attention and a feed-forward network, held together by residual connections and layer normalization, to create rich, contextual representations of an input sequence.
Today, we shift our focus to the other half of the original architecture: the decoder. The decoder's role is to take the encoder's output and generate a new sequence, one token at a time. This generative capability is at the heart of models like GPT. Your learning outcome for this lesson is to build a Transformer decoder block with masked multi-head self-attention.
By the end of this lesson, you will understand:
- The three sub-layers that make up a standard decoder block.
- The critical concept of masked self-attention and why it's essential for generation.
- The function of cross-attention, where the decoder queries the information processed by the encoder.
- The difference between a decoder block in an encoder-decoder model (like for translation) and a decoder-only model (like for generation).
1. The Anatomy of a Transformer Decoder Block
At first glance, the decoder block looks more complex than the encoder block. It has three main sub-layers instead of two.

The three sub-layers are:
- Masked Multi-Head Self-Attention: Similar to the encoder's self-attention, but with a crucial "mask" to prevent it from "cheating" by looking at future tokens.
- Encoder-Decoder Cross-Attention: This is a new component. It's where the decoder pays attention to the output of the encoder.
- Position-wise Feed-Forward Network (FFN): This is identical in structure to the one in the encoder block.
Let's dissect each of these, starting with the most important innovation: the mask.
2. Sub-layer 1: Masked Multi-Head Self-Attention
The primary purpose of a decoder in a generative task is to predict the next token in a sequence. During training, we feed the decoder the entire target sequence. However, to mimic the step-by-step generation process, we must ensure that when predicting the token at position i, the model can only use information from tokens 1 to i. It cannot see tokens i+1, i+2, etc. This is called causality, and we enforce it with a look-ahead mask.
The Masking Mechanism
The mask is typically a matrix that nullifies the attention scores for future positions before they are fed into the softmax function. This ensures that the probability distribution for attention weights only considers current and past tokens.
For an intuitive and elegant explanation of how this is implemented efficiently, let's turn to Andrej Karpathy. He demonstrates a "mathematical trick" using a lower triangular matrix to achieve this masking during a weighted average calculation, which is the core of attention.
Let's build GPT: from scratch, in code, spelled out.
In this segment from 'Let's build GPT', Andrej Karpathy explains the core logic behind causal masking using a toy example. This provides a strong foundation for understanding how masked self-attention works.
Watch the video from 00:42:18 to 00:57:56. Focus on: The goal: allowing tokens to communicate only with their past. The inefficient for loop approach to averaging past information. The key insight: using matrix multiplication with a special matrix (a lower triangular matrix of weights) to perform this 'causal averaging' efficiently.
This trick, where we set the upper-triangular part of the attention score matrix to negative infinity (-inf), is precisely how masked self-attention is implemented. When softmax is applied, these -inf values become zero, effectively nullifying any attention to future tokens.
The following video from CodeEmporium provides a very clear, visual representation of this look-ahead mask.
Transformer Decoder coded from scratch
This short clip clearly illustrates the structure of the look-ahead mask and its purpose in preventing the model from 'cheating' during training.
Watch from 00:06:07 to 00:07:15. Observe the triangular pattern of zeros and negative infinities and connect it to the concept of allowing a token to only see itself and the tokens that came before it.
With this mask applied, the rest of the multi-head self-attention mechanism works exactly as it did in the encoder.
3. Sub-layer 2: Encoder-Decoder Cross-Attention
This is where the decoder gets to "look at" the source sentence that the encoder processed. Without this, the decoder wouldn't know what it's supposed to be translating or responding to.
The mechanism is still multi-head attention, but the inputs for Query (Q), Key (K), and Value (V) come from different places:
- Query (Q): Comes from the output of the previous sub-layer in the decoder (the masked self-attention layer). You can think of this as the decoder asking a question: "Based on the text I've generated so far, what information do I need from the source sentence?"
- Key (K) and Value (V): Both come from the final output of the encoder stack. This represents the full contextualized meaning of the source sentence. The decoder uses its query to match against the keys of the source sentence's tokens to decide which values are most important to focus on.
Crucially, there is no mask in cross-attention. The decoder is allowed to attend to all parts of the source sentence at every step of the generation process.
The CodeEmporium video provides an excellent, detailed walkthrough of implementing this cross-attention mechanism.
Transformer Decoder coded from scratch
This segment explains the implementation of multi-head cross-attention, highlighting the critical difference from self-attention: the origins of the Q, K, and V vectors.
Watch from 00:28:05 to 00:34:19. Pay close attention to how the queries are generated from the decoder's input (y) while the keys and values are generated from the encoder's output (x). This separation is the defining feature of cross-attention.
4. Putting It All Together: A Full Decoder Block
Now we can assemble the full decoder block. The data flow is as follows:
- The input (the target sequence generated so far) goes through Masked Multi-Head Self-Attention.
- A residual connection adds the output to the input, followed by Layer Normalization.
- The result goes to the Cross-Attention layer as the Query, with the Encoder's output providing the Key and Value.
- Another residual connection and Layer Normalization are applied.
- The result passes through the Feed-Forward Network (FFN).
- A final residual connection and Layer Normalization are applied.
The following resource provides a clean PyTorch implementation of this exact structure in a DecoderLayer class.
Transformer Model Tutorial in PyTorch: From Theory to Code
This DataCamp tutorial provides a clear, well-commented implementation of a Transformer decoder block for an encoder-decoder model. It perfectly summarizes the architecture we've just discussed.
Read Section 4, 'Building the decoder blocks'. Study the DecoderLayer class. Trace the forward method and identify the three main sections corresponding to self-attention, cross-attention, and the feed-forward network, along with their respective norm and dropout calls.
5. Important Context: Decoder-Only Transformers (GPT-style)
The block we just built is for an encoder-decoder architecture, like the one in the original "Attention Is All You Need" paper, which was designed for machine translation.
However, many of the most famous modern language models, like GPT, are decoder-only Transformers. They are not conditioned on a separate encoded input; their only job is to continue a sequence of text.
How does a decoder-only block differ? It's much simpler: you just remove the cross-attention sub-layer.
A decoder-only block consists of only:
- Masked Multi-Head Self-Attention
- A Position-wise Feed-Forward Network
This is essentially an encoder block, but with causal masking enabled in its self-attention layer. The article below provides a fantastic, concise implementation of exactly this kind of block, which is the workhorse of generative LLMs.
Decoder-Only Transformers: The Workhorse of Generative LLMs
This article from 'Decoder-Only Transformers: The Workhorse of Generative LLMs' shows the implementation of a decoder-only block. Comparing this code to the DecoderLayer from the previous resource will make the distinction crystal clear.
Read the section 'Putting It All Together!'. Examine the Block class. Notice its simplicity: it only contains self.ln_1, self.attn, self.ln_2, and self.ffnn. There is no cross-attention component.
This distinction is crucial for understanding the landscape of modern AI architectures.
Test your understanding!
- Why is masking necessary in the decoder's self-attention layer but not in its cross-attention layer?
- In an encoder-decoder architecture, you are translating a sentence. During the cross-attention step in the third decoder block, where do the Q, K, and V vectors originate from?
- You are given the
DecoderLayerclass from the DataCamp tutorial, which includes cross-attention. How would you modify it to create aDecoderOnlyBlockfor a GPT-style generative model?
Show answer
- Masking is necessary in self-attention to maintain the autoregressive property. When generating the
i-th token, the decoder should only know about tokens1toiof the sequence it is generating. Masking prevents it from "cheating" by looking at future tokens. Masking is not needed in cross-attention because the decoder should have access to the entire source sentence (from the encoder) at every step to make an informed translation. -
- Query (Q): Comes from the output of the masked self-attention sub-layer within the third decoder block.
- Key (K) and Value (V): Both come from the final output of the entire encoder stack (i.e., after the source sentence has passed through all N encoder blocks). The K and V from the encoder are the same for every decoder block.
- You would remove the cross-attention sub-layer and its associated layer normalization and dropout. Specifically, you would delete
self.cross_attnandself.norm2from the__init__method, and remove the entire cross-attention section from theforwardmethod:
You would also rename# This entire block would be removed attn_output = self.cross_attn(x, enc_output, enc_output, src_mask) x = self.norm2(x + self.dropout(attn_output))norm1andnorm3toln_1andln_2for clarity, and remove theenc_outputandsrc_maskarguments from theforwardmethod's signature.
Conclusion
You've now dissected the Transformer decoder block, a critical component for all generative sequence models.
Key Takeaways:
- A standard decoder block has three sub-layers: Masked Self-Attention, Cross-Attention, and a Feed-Forward Network.
- Masked Self-Attention enforces causality, preventing the model from looking at future tokens in the sequence it is generating.
- Cross-Attention is the communication channel where the decoder queries the final representation of the source sequence from the encoder.
- Decoder-Only Transformers (like GPT) use a simplified block that omits the cross-attention layer, as they are not conditioned on a separate encoded input.
Preview of the Next Lesson:
We've now built both the encoder and decoder blocks. You've seen "Add & Norm" applied after every sub-layer. In the next lesson, we will consolidate our understanding of these two crucial techniques: layer normalization and residual connections. We'll review their importance for training stability, compare the common "pre-norm" vs. "post-norm" formulations, and ensure we have a solid grasp of how they are applied consistently across the entire Transformer architecture before we move on to assembling the full model.