Hello! Welcome to the final lesson in our module on the Transformer architecture. Over the past few lessons, we have meticulously built and examined each of the fundamental components:
- Positional Encodings
- Multi-Head Self-Attention and Masked Self-Attention
- Position-wise Feed-Forward Networks
- Encoder and Decoder Blocks
- The crucial "Add & Norm" layers with Residual Connections and Layer Normalization
Today, we'll put all the pieces together. Your learning outcome for this lesson is to assemble a full encoder-decoder Transformer model from scratch. We will see how the encoder and decoder stacks interact to form a powerful sequence-to-sequence model, capable of tasks like machine translation.
By the end of this lesson, you will understand:
- The overall data flow through a complete Transformer model.
- How to construct the full
EncoderandDecodermodules by stacking their respective blocks. - How the
Encoder's output provides context to theDecodervia cross-attention. - How to combine these modules into a final, top-level
Transformerclass. - The difference between the encoder-decoder architecture and decoder-only architectures like GPT.
1. The Big Picture: Encoder-Decoder Architecture
The original Transformer, as described in "Attention Is All You Need," is an encoder-decoder model. This design is ideal for sequence-to-sequence (seq2seq) tasks, where you transform an input sequence (e.g., a sentence in English) into a different output sequence (e.g., the same sentence in German).
The architecture consists of two main parts:
- The Encoder Stack: Processes the entire input sequence and generates a rich, contextual representation of it.
- The Decoder Stack: Receives the encoder's output and the beginning of the target sequence, then autoregressively generates the rest of the target sequence, one token at a time.
Let's look at the canonical diagram of the full model.

The process works as follows:
- Encoder Phase: The source sentence (e.g., "I love you") is tokenized, embedded, and passed through the stack of N encoder layers. The final output is a sequence of vectors (often called
memoryorencoder_output) that represents the source sentence. - Decoder Phase:
- The decoder is given the start of the target sentence (initially just a special
<start-of-sentence>token). - This input is tokenized, embedded, and passed through the stack of N decoder layers.
- In each decoder layer, the model first uses masked self-attention to consider the previously generated target tokens.
- Then, it uses cross-attention to look at the
encoder_output, allowing it to draw context from the source sentence. - The final output of the decoder stack is passed through a linear layer and a softmax function to predict the next token in the target sequence.
- This new token is added to the decoder's input, and the process repeats until an
<end-of-sentence>token is generated.
- The decoder is given the start of the target sentence (initially just a special
To see this entire process in motion, let's watch a detailed walkthrough of the training and inference phases.
Attention is all you need (Transformer) - Model explanation (including math), Inference and Training
The video 'Attention is all you need (Transformer)' by Umar Jamil provides an excellent animated explanation of how the encoder and decoder work together during both training and inference. This will help you visualize the complete data flow.
Please watch the following sections: Overall Structure (07:56 - 08:48): A quick overview of how the encoder and decoder macro-blocks are connected. Training (44:43 - 52:05): Pay close attention to how the complete source and target sentences are fed into the model simultaneously during training, and how the loss is calculated. Inference (52:05 - 57:06): Focus on the autoregressive, token-by-token generation process, where the output of one step becomes the input for the next.
2. Assembling the Transformer in Code
Given your background as a software engineer, the most intuitive way to understand the assembly is to think in terms of object composition. We have our building-block modules (MultiHeadAttention, PositionwiseFeedForward, EncoderBlock, DecoderBlock). Now, we'll compose them into higher-level modules: Encoder, Decoder, and finally, Transformer.
We will follow the structure of a well-regarded PyTorch implementation. The linked repository provides a clear, modular implementation that is easy to follow.
A PyTorch Tutorial to Transformers
The GitHub repository 'a-PyTorch-Tutorial-to-Transformers' by Sagar Vinod is an excellent, code-first guide. We will use it to understand how to structure the full model in PyTorch.
Read the sections explaining the Encoder, Decoder, and Transformer classes. You can refer to the actual model.py file linked in the article to see the full code. Focus on the __init__ and forward methods of each class: Encoder (model.py): Notice how it takes the input embeddings, adds positional encodings, and then passes the result through a nn.ModuleList of EncoderLayers. Decoder (model.py): See how its forward method takes not only the decoder input (dec_sequences) but also the enc_sequences (the output from the Encoder). This enc_sequences is passed to the cross-attention sublayer within each DecoderLayer. Transformer (model.py): Observe how this top-level class simply instantiates the Encoder and Decoder. Its forward method orchestrates the entire process: call the encoder, then call the decoder with the encoder's output.
Let's summarize the key points of this code structure.
The Encoder Module
The Encoder is responsible for processing the source sequence.
__init__:- Takes hyperparameters like vocabulary size, model dimension (
d_model), number of layers (N), number of heads (h), dropout rate, etc. - Creates an embedding layer for source tokens.
- Creates a positional encoding layer.
- Creates a
nn.ModuleListcontainingNinstances of theEncoderBlockyou built previously. - Creates a final
LayerNorm.
- Takes hyperparameters like vocabulary size, model dimension (
forward(src_tokens):- Converts
src_tokensto embeddings. - Adds positional encodings.
- Passes the result through the stack of
EncoderBlocks in a loop. - Applies the final layer normalization.
- Returns the final representation (
encoder_output).
- Converts
The Decoder Module
The Decoder is responsible for generating the target sequence.
__init__:- Similar to the encoder, it creates embedding layers, positional encodings, and a
nn.ModuleListofNDecoderBlocks. - It also creates the final linear "classification head" that projects the output vector to the vocabulary size.
- Similar to the encoder, it creates embedding layers, positional encodings, and a
forward(tgt_tokens, encoder_output):- This is the crucial part. It takes two main arguments.
- Converts
tgt_tokensto embeddings and adds positional encodings. - Passes the result through the stack of
DecoderBlocks. Crucially, eachDecoderBlockreceives theencoder_outputto perform cross-attention. - Passes the final output through the linear classification head to get logits.
- Returns the logits.
The Top-Level Transformer Module
This class ties everything together.
__init__:- Instantiates the
Encodermodule. - Instantiates the
Decodermodule. - May handle weight initialization and tying the weights of the embedding layers and the final classification head (a common practice).
- Instantiates the
forward(src_tokens, tgt_tokens):encoder_output = self.encoder(src_tokens)logits = self.decoder(tgt_tokens, encoder_output)return logits
This clean, hierarchical structure is a hallmark of good software engineering and is exactly how complex deep learning models are built in practice.
Test your understanding!
Imagine you are debugging the forward pass of a full Transformer model designed for English-to-German translation.
- What are the dimensions and meaning of the
encoder_outputtensor that is passed from the encoder to the decoder? - Inside a
DecoderBlock, how many times is theencoder_outputtensor used, and for what purpose? - Why is it essential that the
Decoder's self-attention mechanism is masked, while theEncoder's is not?
Show answer
- The
encoder_outputtensor has dimensions(batch_size, source_sequence_length, d_model). It is a sequence of vectors, where each vector is a rich, contextual representation of the corresponding token in the source (English) sentence. It contains the "meaning" of the entire source sentence, which the decoder will consult during translation. - The
encoder_outputis used once inside eachDecoderBlock. It serves as the Key (K) and Value (V) for the cross-attention (or encoder-decoder attention) sub-layer. The Query (Q) comes from the output of the preceding masked self-attention layer. This allows each token in the target (German) sequence to "look at" all tokens in the source (English) sequence to decide what information is most relevant for predicting the next German word. - The masking is critical for preserving the autoregressive property of the decoder. During training, the decoder is fed the entire ground-truth target sequence at once for efficiency. The mask prevents a token at position
ifrom "cheating" by looking at future tokens at positionsi+1,i+2, etc. It can only attend to past tokens (and itself). The encoder, however, processes a static source sentence that is fully known. It is beneficial for it to build the best possible representation, so every token is allowed to attend to every other token in the input.
3. A Quick Detour: Encoder-Decoder vs. Decoder-Only
The architecture we just built is the "full" Transformer. However, many famous models you've heard of, like GPT, are decoder-only models. What's the difference?
- Encoder-Decoder (e.g., original Transformer, T5, BART): Best for transforming one sequence into another (translation, summarization). They have both an encoder to understand the source and a decoder to generate the target. They contain self-attention and cross-attention.
- Decoder-Only (e.g., GPT series, LLaMA): Best for pure generation or completing a prompt. They are essentially just the decoder stack. They take a sequence and predict the next token. They only have masked self-attention, as there is no "source" sequence to cross-attend to.
Andrej Karpathy gives a fantastic, succinct explanation of this distinction.
Let's build GPT: from scratch, in code, spelled out.
In his 'Let's build GPT' tutorial, Andrej Karpathy explains exactly why the original paper has an encoder-decoder structure (for machine translation) and why models like GPT are decoder-only.
Watch from 01:42:37 to 01:46:15. This will perfectly contextualize the model we just assembled and contrast it with the generative models we will study in more detail later.
Conclusion
Congratulations! You have now assembled a complete Transformer from all the constituent parts we've studied. You've gone from the smallest details of dot-product attention all the way up to a full encoder-decoder architecture.
Key Takeaways:
- The full Transformer model consists of an Encoder stack and a Decoder stack.
- The Encoder processes the source sequence and passes its final representation (
encoder_output) to the Decoder. - The Decoder uses this
encoder_outputin its cross-attention layers to condition its generation on the source sequence. - The entire model can be built in a modular, object-oriented way, composing blocks into stacks and stacks into the final model.
- This encoder-decoder architecture is designed for seq2seq tasks, distinct from decoder-only models (like GPT) which are for pure generation.
Preview of the Next Module:
We've built the engine. Now, we need to talk about the fuel. How do we convert raw text into the numerical sequences of tokens and embeddings that this model understands? The next module, Foundations of Language Modeling and Embeddings, will kick off by exploring Word2Vec, a foundational technique for learning meaningful vector representations of words.