Hello! Welcome to the fourth lesson in our module on Sequence Modeling.
In our last lesson, we took a deep dive into the self-attention mechanism, the engine that powers the Transformer. We saw how it uses Queries, Keys, and Values to create context-aware representations, and we highlighted its key advantages over RNNs: direct, path lengths between any two tokens and massive parallelizability.
However, we also left a loose end: self-attention, by itself, is permutation invariant—it has no inherent sense of word order. We also only looked at it in isolation. Today, we'll address that and assemble the complete puzzle.
The learning outcome for this lesson is to: Describe the complete Transformer architecture, including positional encoding, multi-head attention, and feed-forward layers. We will construct the model brick by brick, from the input embeddings to the final output probabilities, understanding how each component contributes to the whole.
1. The Transformer: A High-Level View
At its core, the Transformer maintains the familiar encoder-decoder structure common in sequence-to-sequence tasks like machine translation or speech recognition.

- The Encoder Stack: Its job is to process the entire input sequence (e.g., an audio spectrogram or a sentence) and build a rich, context-aware representation of it.
- The Decoder Stack: Its job is to generate the output sequence (e.g., a transcript) one element at a time. It uses the encoder's representation to inform its generation process.
Unlike earlier models that used RNNs for the encoder and decoder, the Transformer uses stacks of identical layers built around self-attention. The original paper used a stack of 6 encoder layers and 6 decoder layers.
Let's start by looking at the components that make up a single encoder layer.
For a clear, high-level overview, let's start with a resource that excels at visual explanations. Jay Alammar's "The Illustrated Transformer" provides an intuitive entry point into the model's structure.
Read the first two sections, "A High-Level Look" and "Bringing The Tensors Into The Picture". This will help you visualize the overall encoder-decoder stack and the two main sub-layers within each encoder: self-attention and a feed-forward network.
2. Building an Encoder Layer
An encoder layer must process a sequence of input vectors and produce a sequence of output vectors of the same dimension, incorporating contextual information. It does this using four key components.
2.1. Positional Encoding
As we discussed, the self-attention mechanism is permutation-invariant. To give the model information about the order of the sequence, we inject a "positional encoding" vector into each input embedding.
These vectors are not learned; they are calculated using a clever function involving sine and cosine waves of different frequencies.
Where:
posis the position of the token in the sequence.iis the dimension within the embedding vector.- is the dimension of the embedding (e.g., 512).
The intuition here is that for any fixed offset , can be represented as a linear function of . This property is thought to make it easy for the model to learn to attend to relative positions. The positional encoding vector is simply added to the input embedding.
Attention is all you need (Transformer) - Model explanation (including math), Inference and Training
Let's watch a detailed explanation of positional encoding, covering its motivation and the mathematics behind it.
Watch the segment from 15:04 to 20:08. Pay attention to how the sine and cosine functions create a unique positional signature for each token that the model can learn to interpret.
2.2. Multi-Head Attention
In the last lesson, we learned about single-head self-attention. The Transformer improves on this with Multi-Head Attention. Instead of calculating attention once, it does it multiple times in parallel.
This has two main benefits:
- Focus on Different Relationships: It allows the model to jointly attend to information from different "representation subspaces." For example, one attention head might learn to track syntactic dependencies, while another tracks semantic relationships.
- Expanded Focus: It helps the model focus on multiple relevant words at once. In the sentence, "The animal didn't cross the street because it was too tired," one head might learn for "it" to attend to "animal," while another head learns for "it" to attend to "tired."

The mechanism works as follows:
- Projection: The input vectors (or the Q, K, V matrices) are not fed into one attention function. Instead, they are linearly projected times (where is the number of heads) using different, learned weight matrices ( for each head ). This creates sets of queries, keys, and values.
- Parallel Attention: Scaled dot-product attention is applied in parallel to each of these projected sets, yielding output matrices.
- Concatenation & Final Projection: The output matrices are concatenated and then projected back to the original model dimension with another learned weight matrix, .
Jay Alammar's post provides an excellent visual breakdown of this process.
Read the section "The Beast With Many Heads". Focus on how the multiple Q/K/V sets are created and how their outputs are combined.
2.3. Position-wise Feed-Forward Network (FFN)
After the multi-head attention sub-layer, the output for each position is passed through an identical but separate Feed-Forward Network (FFN). This component adds further processing capacity and non-linearity.
It consists of two linear transformations with a ReLU activation in between:
- The input and output dimensions are .
- The inner-layer dimension is typically larger (e.g., ). For a 512-dimension model, this would be .
This network is called "position-wise" because it processes each position's vector independently.
2.4. Residuals and Layer Normalization
To make the network deep and stable, two more crucial elements are added. Each of the two sub-layers (Multi-Head Attention and FFN) in an encoder layer has a residual connection around it, followed by layer normalization.
The output of each sub-layer is calculated as: .
- Residual Connection: This is the
x + ...part, a concept from ResNets. It helps combat the vanishing gradient problem in deep networks and allows for easier information flow. - Layer Normalization: This normalizes the outputs of the sub-layer, stabilizing the training dynamics. Unlike batch normalization, which normalizes across the batch, layer normalization normalizes across the features for each individual sequence.
With these four components, we can now define a full encoder layer.
3. Assembling the Decoder
The decoder's job is to generate the output sequence. It shares many components with the encoder, but with a few key differences to handle its generative, auto-regressive nature. A decoder layer has three sub-layers instead of two.
-
Masked Multi-Head Self-Attention: The decoder first performs self-attention on the output sequence it has generated so far. However, to maintain the auto-regressive property (i.e., the prediction for the current token can only depend on previous tokens), we must prevent it from "peeking ahead." This is done by applying a mask to the attention scores matrix before the softmax step. The mask sets the scores for all future positions to , which makes their softmax probabilities zero.
-
Encoder-Decoder Attention: This is where the decoder interacts with the encoder. This sub-layer is a multi-head attention mechanism, but with a twist:
- The Queries (Q) come from the output of the previous decoder layer (the masked self-attention).
- The Keys (K) and Values (V) come from the output of the final encoder layer.
This allows every position in the decoder to attend to all positions in the input sequence, helping it decide which parts of the input are most relevant for generating the next output token.
-
Position-wise Feed-Forward Network: This is identical to the FFN in the encoder.
Like the encoder, each of these three sub-layers is also wrapped in a residual connection and followed by layer normalization.
4. The Full Picture: Training and Inference
To put it all together, let's trace the data flow during training.
Attention is all you need (Transformer) - Model explanation (including math), Inference and Training
This video segment walks through the entire Transformer architecture, showing how the encoder and decoder interact during training. It visually connects all the components we've discussed.
Watch the segment from 44:43 to 52:14. This part is crucial for understanding how the full system works end-to-end in a single training pass.
To summarize the training process:
- Encoder Pass: The entire source sequence is fed into the encoder stack. The output of the top encoder layer (the K and V matrices) is produced in one go.
- Decoder Pass: The entire target sequence, shifted right by one position and prepended with a
start-of-sequencetoken, is fed into the decoder stack. - Masked Self-Attention: The decoder uses masked self-attention to process the target sequence without looking ahead.
- Encoder-Decoder Attention: At each decoder layer, the model uses the encoder's output (K and V) to cross-reference the source sequence.
- Final Output Layer: The output from the top decoder layer is passed through a final linear layer (to project to the vocabulary size) and a softmax function to get the probability distribution over the next possible token.
- Loss Calculation: The output probabilities are compared against the actual target sequence (the ground truth), and a loss (e.g., cross-entropy) is calculated to update the model's weights via backpropagation.
Crucially, all of this happens in a single, parallelizable forward pass, which is why Transformers are so efficient to train compared to RNNs.
5. From Theory to Code
Your background in computer science and experience with PyTorch will make it valuable to see how these theoretical blocks translate into code. While we will implement these components in later lessons, watching a guided implementation can solidify your understanding now.
Pytorch Transformers from Scratch (Attention is all you need)
The following video by Aladdin Persson provides a from-scratch implementation of the Transformer in PyTorch. Watching it will help you connect the abstract diagrams and formulas to concrete nn.Module classes.
You don't need to code along, but watch the following sections to see how the architecture is built: Transformer Block (27:01 - 32:18): See how SelfAttention, LayerNorm, and the FFN are combined into a reusable block. Encoder (32:18 - 38:21): Watch how the TransformerBlock is stacked to create the full encoder, including the embedding and positional encoding. Decoder Block & Decoder (38:21 - 46:52): Observe the implementation of the decoder's three sub-layers and how they are stacked. Full Transformer (46:52 - 52:38): See the final model assembly, bringing the encoder and decoder together and defining the masking logic. \nThis will give you a strong mental model of the practical implementation.
Conclusion
In this lesson, we assembled the complete Transformer architecture from its fundamental building blocks. We've seen how it moves beyond the simple self-attention mechanism to create a deep, powerful, and highly parallelizable model for sequence-to-sequence tasks.
Key Takeaways:
- The Transformer consists of an encoder stack and a decoder stack.
- Positional Encodings are added to the input embeddings to provide the model with sequence order information.
- Multi-Head Attention allows the model to learn different types of relationships in parallel by using multiple sets of Q, K, V projections.
- Each layer contains Feed-Forward Networks for additional non-linear processing and is wrapped with residual connections and layer normalization for stable training.
- The decoder has three sub-layers: a masked self-attention layer to preserve auto-regression, an encoder-decoder attention layer to consult the input sequence, and an FFN.
- The entire model can be trained in a single, parallel forward-backward pass, a major advantage over recurrent models.
Preview of the Next Lesson:
Now that we have a solid theoretical and high-level code understanding of the Transformer's architecture, our next step is to get our hands dirty. In the next lesson, we will begin the practical implementation, starting with the core of the model: implementing a multi-head self-attention layer in PyTorch, including support for causal masking.