Skip to main content
Create your own

Building a Transformer Encoder Block

Hello! Welcome back to our deep dive into the Transformer architecture.

In our previous lessons, we've laid all the necessary groundwork. We started by implementing the multi-head self-attention mechanism, which allows the model to understand the relationships between different tokens in a sequence. Then, we solved the permutation-invariance problem by implementing sinusoidal positional encodings, giving the model a sense of order.

Now, it's time to assemble these components into a functional unit. Your learning outcome for this lesson is to build a complete Transformer encoder block, including multi-head attention and a position-wise feed-forward network. This block is the fundamental, repeatable building block of the entire encoder side of the Transformer.

1. Anatomy of the Transformer Encoder Block

The encoder's job is to take the sequence of input embeddings (with positional information added) and transform it into a sequence of rich, context-aware representations. It doesn't do this in one go. Instead, it uses a stack of identical "encoder blocks" to progressively refine these representations.

Let's look at the structure of a single block.

Transformer Encoder Block Architecture
This diagram shows the two main sub-layers of an encoder block: Multi-Head Attention and a Feed-Forward Network. Crucially, it also shows the 'Add & Norm' components, which represent residual connections and layer normalization.

As you can see, each encoder block consists of two main sub-layers:

  1. A Multi-Head Self-Attention layer.
  2. A Position-wise Fully Connected Feed-Forward Network (FFN).

These sub-layers are connected using two other critical components that are essential for training deep networks:

  • Residual Connections (the "Add" part).
  • Layer Normalization (the "Norm" part).

Let's break down how these pieces fit together.

2. The Sub-layers: Attention and Computation

First, let's formally define the components of the block. The "Annotated Transformer" provides the canonical description based on the original paper.

The Annotated Transformer

This section of 'The Annotated Transformer' by Harvard NLP precisely describes the structure of the encoder. It will serve as our formal guide.

Read the subsection titled 'Encoder'. Pay close attention to the description of its two sub-layers and the mention of residual connections and layer normalization. The formula LayerNorm(x + Sublayer(x)) is the key to understanding the data flow.

2.1. Sub-layer 1: Multi-Head Self-Attention

You are already familiar with this from a previous lesson. In the encoder, the input sequence is fed into the multi-head self-attention layer as the Query (Q), Key (K), and Value (V). This allows every token in the input sequence to attend to every other token, generating an updated representation for each token that is now aware of its context. Since all input tokens are processed simultaneously in the encoder, there is no need to mask future tokens as we will see in the decoder.

2.2. Sub-layer 2: Position-wise Feed-Forward Network (FFN)

After the attention mechanism has allowed tokens to exchange and aggregate information, the FFN provides further processing.

It's a simple two-layer Multi-Layer Perceptron (MLP) that is applied independently to each token's representation.

  • Structure: It consists of two linear transformations with a ReLU activation in between.
  • Function: You can think of this as giving the model more computational depth. After a token has gathered context from its neighbors via self-attention, the FFN allows it to "think" on this new information and perform a more complex, non-linear transformation of its own representation.
  • Dimensionality: A common practice, as noted in the original paper, is to make the inner dimension of the FFN larger than the model's dimension (e.g., d_model=512, d_ff=2048). This temporarily expands the representation space, allowing the model to learn more complex patterns before projecting it back to the original d_model.

This tutorial gives a clear explanation of both the Multi-Head Attention and the FFN within the encoder block.

Tutorial 5: Transformers and Multi-Head Attention

The PyTorch Lightning tutorial on Transformers provides excellent, high-level explanations for each part of the encoder block.

Read the section 'Transformer Encoder'. It clearly explains the role of the multi-head attention layer, the residual connections, layer normalization, and the position-wise FFN.

3. The Glue: Residuals and Normalization

Having powerful sub-layers isn't enough; we need to connect them in a way that allows for stable training, especially as we stack many of these blocks to create a deep network.

3.1. Residual Connections

The "Add" part of "Add & Norm" refers to a residual connection. The output of a sub-layer is not just the transformed input Sublayer(x), but rather x + Sublayer(x).

This is a powerful technique borrowed from computer vision (e.g., ResNet). Given your background, you can think of it as creating a "gradient superhighway".

  • During backpropagation, the gradient can flow directly through the + operation, bypassing the sub-layer entirely. This prevents the vanishing gradient problem in very deep networks.
  • It reframes the learning problem: the sub-layer doesn't need to learn the entire, complex transformation from scratch. Instead, it only needs to learn the residual, or the difference, that needs to be added to the original input x.

3.2. Layer Normalization

The "Norm" part is Layer Normalization. It is applied to the output of the residual connection. Its purpose is to stabilize the training process by ensuring the inputs to the next layer have a consistent distribution (e.g., zero mean and unit variance).

It normalizes the features across the embedding dimension for each token in the sequence independently. This is distinct from Batch Normalization, which normalizes across the batch dimension for each feature. LayerNorm is generally preferred in NLP because it's independent of the batch size and works well with variable sequence lengths.

Pre-Norm vs. Post-Norm:
The original paper applied normalization after the residual addition (LayerNorm(x + Sublayer(x))), known as post-norm. However, many modern implementations use a pre-norm formulation (x + Sublayer(LayerNorm(x))), where normalization is applied to the input before it enters the sub-layer. This is often found to be more stable during training. We will see this in the code implementations.

Andrej Karpathy provides a fantastic and intuitive explanation for both residual connections and layer normalization.

Let's build GPT: from scratch, in code, spelled out.

In his 'Let's build GPT' video, Andrej Karpathy explains the importance of residual connections and layer normalization for training deep networks. He also implements the more modern 'pre-norm' formulation.

Watch the segment from 01:26:46 to 01:37:15. Focus on the intuition he provides: Residual connections as a 'gradient superhighway'. Layer Normalization as a way to control the statistics of the activations, similar to BatchNorm but applied differently. The pre-norm formulation where LayerNorm is applied before the sub-layer.

4. Implementation: Building the Encoder

Now, let's translate this architecture into PyTorch code. We will build an EncoderLayer (or EncoderBlock) class that encapsulates all the logic, and then a main Encoder class that stacks these layers.

The following video is a complete, line-by-line walkthrough of building a Transformer Encoder. It's the perfect guide for this section.

Transformer Encoder in 100 lines of code!

This video from CodeEmporium is an end-to-end guide for building the Transformer Encoder in PyTorch. It's concise and perfectly aligned with our learning outcome.

Watch the video in segments, following the code construction. I'll break it down for you: EncoderLayer Class (14:30 - 17:37): This is the core of our lesson. Pay close attention to how the forward method implements the architecture we've discussed: residual, attention, add & norm, residual, FFN, add & norm. Multi-Head Attention (17:37 - 36:36): This is a detailed review of what you've learned before, but now placed in the context of the full encoder. It's worth watching to see how the dimensions (batch_size, seq_len, d_model) flow through the attention heads. Layer Normalization (36:36 - 43:06): A fantastic deep dive into the implementation of LayerNorm. Your CS background will help you appreciate how the mean and standard deviation are calculated over the last dimension (the embedding dimension). Feed-Forward Network (43:06 - 47:12): See the simple but effective two-layer MLP being implemented, including the expansion and contraction of the feature dimension. The Full Encoder (10:42 - 14:30): Finally, see how an Encoder class is built to stack multiple EncoderLayer instances using nn.Sequential.

Let's summarize the key code structure for the EncoderBlock class you saw being built. Note that this implementation follows the post-norm structure from the original paper, which is a great starting point.

import torch.nn as nn
from multi_head_attention import MultiHeadAttention # Assuming you have this from a previous lesson
from feed_forward import PositionwiseFeedForward # A class for the FFN

class EncoderBlock(nn.Module):
    def __init__(self, d_model, num_heads, d_ff, dropout=0.1):
        super(EncoderBlock, self).__init__()
        
        # Self-Attention sub-layer
        self.self_attn = MultiHeadAttention(d_model, num_heads)
        self.norm1 = nn.LayerNorm(d_model)
        self.dropout1 = nn.Dropout(dropout)
        
        # Feed-Forward sub-layer
        self.ffn = PositionwiseFeedForward(d_model, d_ff)
        self.norm2 = nn.LayerNorm(d_model)
        self.dropout2 = nn.Dropout(dropout)

    def forward(self, x, mask=None):
        # Attention sub-layer
        # Note: In Karpathy's video you saw pre-norm: x_norm = self.norm1(x)
        attn_output, _ = self.self_attn(x, x, x, mask) 
        x = self.norm1(x + self.dropout1(attn_output)) # Post-norm: Add, then Norm
        
        # Feed-Forward sub-layer
        # Note: Pre-norm would be: ffn_output = self.ffn(self.norm2(x))
        ffn_output = self.ffn(x)
        x = self.norm2(x + self.dropout2(ffn_output)) # Post-norm: Add, then Norm
        
        return x
Test your understanding!
  1. What is the purpose of the residual connection in the Transformer block? Why not just pass the output of one sub-layer directly to the next?
  2. The Position-wise Feed-Forward Network is applied "identically" to each position. If you have an input tensor of shape (batch_size=32, seq_len=100, d_model=512), do the 100 tokens in a sequence share the weights of the FFN? Do the 32 sequences in the batch share the weights?
  3. Why is Layer Normalization generally preferred over Batch Normalization in Transformer models for NLP?
Show answer
  1. The residual connection x + Sublayer(x) helps mitigate the vanishing gradient problem in deep networks by creating a "shortcut" for gradients to flow backward to the input. This makes training deep stacks of Transformer blocks much more stable.
  2. Yes, the weights of the FFN are shared across all 100 token positions within a sequence. This is what "position-wise" and "identically" means. The same network is applied to each token's embedding. Yes, the weights are also shared across all 32 sequences in the batch. The linear layers in the FFN have a single set of weights that are used for every token of every sequence in the batch.
  3. Batch Normalization calculates statistics across the batch dimension. In NLP, batch sizes can be small due to memory constraints, and sequence lengths can vary, which can lead to noisy and unstable statistics for BatchNorm. Layer Normalization calculates statistics across the feature/embedding dimension for each sequence element independently, making it robust to variations in batch size and sequence length.

Conclusion

Congratulations! You have now assembled the complete Transformer encoder block. This powerful and repeatable unit is the workhorse of many state-of-the-art AI models, from BERT to the Vision Transformer (ViT).

Key Takeaways:

  • A Transformer Encoder block is composed of two sub-layers: Multi-Head Self-Attention and a Position-wise Feed-Forward Network.
  • Residual Connections (x + Sublayer(x)) are used around each sub-layer to enable deep stacking and stable gradient flow.
  • Layer Normalization is applied after each residual connection to stabilize activations and speed up training. Modern implementations often use a "pre-norm" variation for even better stability.
  • The full encoder is created by stacking these blocks N times, with the output of one block feeding into the next.

Preview of the Next Lesson:
We have now fully explored the encoder side of the original "Attention Is All You Need" architecture. Next, we will turn our attention to the other half. In the upcoming lesson, "Build a Transformer decoder block with masked multi-head self-attention," we will see how the decoder differs. You will learn about:

  1. Masked Self-Attention, which prevents a position from attending to future positions, a crucial feature for generating sequences one token at a time.
  2. A second Cross-Attention layer, which allows the decoder to look at the final output representations from the encoder stack.

Can't find a good explanation? Sign up and we'll make it for you

Sign up