Skip to main content
Create your own

GPT Architecture and Causal Attention

Hello! Welcome back to our module on Modern Language Model Architectures.

In our last lesson, we explored the architecture of BERT, an encoder-only model designed for deep language understanding. We saw how its bidirectional self-attention allows it to build representations based on the full context of a sentence, making it a powerhouse for analysis tasks like classification and question answering.

Today, we shift our focus to the other titan of the Transformer world: the Generative Pre-trained Transformer (GPT). Where BERT is designed to understand, GPT is designed to create. This lesson will dissect the architecture that powers models like GPT-3 and ChatGPT.

Our learning outcome is to analyze the architecture of GPT models and the role of causal attention. We will explore why its decoder-only structure is perfectly suited for text generation and dive deep into the mechanism that enables it to write coherent, creative, and contextually relevant text, one token at a time.

1. The GPT Blueprint: A Decoder-Only Autoregressive Model

The core architectural decision behind GPT is to use only the decoder part of the original Transformer. In the last lesson, we saw how BERT discarded the decoder to focus on encoding. GPT does the opposite: it discards the encoder to focus entirely on decoding, or generation.

GPT Model Architecture Diagram
This diagram shows the complete architecture of a GPT model. On the left, you see the overall structure: input embeddings are fed into a stack of identical Transformer Blocks. The output is then processed by a final linear layer to predict the next token. On the right is a detailed view of a single block, which we will dissect throughout this lesson.

The fundamental task of a GPT model is autoregressive language modeling: given a sequence of tokens, predict the very next token. By repeating this process—feeding its own prediction back into the input—the model can generate text of any length.

Let's build GPT: from scratch, in code, spelled out.

To understand why a 'decoder-only' architecture is used and how it differs from the models we've seen before, let's watch a segment from Andrej Karpathy's fantastic 'Let's build GPT' video. He clearly explains the distinction between encoder-decoder models (used for tasks like translation) and decoder-only models (used for pure generation).

Watch from 01:42:26 to 01:46:15. Focus on the core idea: since we are not conditioning our generation on another piece of text (like translating from French to English), the encoder is unnecessary. We only need the generative 'decoder' part, which uses a special kind of attention.

This decoder-only design, built for next-token prediction, has become the de-facto standard for nearly all modern large language models (LLMs) used for generation.

Decoder-Only Transformers: The Workhorse of Generative LLMs

This article from Cameron R. Wolfe provides an excellent textual overview of the architecture. Let's start with the introduction and the section that lays out the full model structure.

Read the introductory section ('The current pace of AI research...') and the section titled 'The Full Decoder-Only Transformer'. These will give you a solid high-level framework of the model's components, from input construction to the final output classification head.

2. Causal Attention: The Secret to Generation

The defining feature of GPT's decoder block is masked self-attention, also known as causal self-attention. The word "causal" here is key: it means a token's representation can only be influenced by the tokens that came before it (the "causes"). It is strictly forbidden from "seeing" future tokens.

This is the fundamental difference from BERT's bidirectional attention. BERT can see the whole sentence to understand a word's meaning. GPT, to predict the next word, must only use the context up to the current position. If it could see the future, the prediction task would be trivial and it would never learn to generate text.

Attention in transformers, step-by-step | Deep Learning Chapter 6

For a crystal-clear conceptual explanation of how this masking works, let's turn to this video from 3Blue1Brown.

Watch from 11:09 to 12:45. The video explains why you can't allow later words to influence earlier ones during training and how this is implemented by 'masking' — setting the attention scores for future tokens to negative infinity before the softmax step.

How Causal Attention is Implemented

So, we need a mechanism where each token i can attend to all tokens from 0 to i, but not to i+1 and beyond. As the 3Blue1Brown video mentioned, this is done by modifying the attention score matrix before the softmax function is applied.

Let's break down the full self-attention process in a GPT decoder block:

  1. Project to Q, K, V: For each token in the input sequence, we create a Query (Q), a Key (K), and a Value (V) vector through linear projections.

    • Query: "What am I looking for?"
    • Key: "What information do I contain?"
    • Value: "What information will I provide if I'm attended to?"
  2. Calculate Attention Scores: We compute the dot product between the Query vector of the current token and the Key vectors of all tokens in the sequence (including itself). This results in a matrix of raw attention scores: Scores = Q @ K^T. This score indicates the relevance of each token to the current one.

  3. Apply the Causal Mask: This is the crucial step. We create a mask, which is a lower-triangular matrix. We add this mask to the score matrix, which effectively sets all scores for future tokens to -infinity.

    • masked_Scores = Scores + mask (where the mask has 0 for positions we keep and -inf for positions we want to discard).
  4. Normalize with Softmax: We apply a softmax function across the rows of the masked_Scores matrix. The -inf values become 0 after softmax, ensuring that future tokens have zero influence. The remaining scores are normalized to sum to 1.

    • Attention_Weights = softmax(scaled(masked_Scores)) (scaled by 1/sqrt(head_dimension))
  5. Aggregate Values: Finally, we compute the new representation for the current token by taking a weighted sum of all the Value vectors, using the Attention_Weights as the weights.

    • Output = Attention_Weights @ V

Decoder-Only Transformers: The Workhorse of Generative LLMs

Now, let's read a textual explanation that walks through this entire process, including the multi-head aspect, and shows the corresponding PyTorch code. Given your background, seeing the implementation will connect the theory to practice.

Read the sections 'Causal Self-Attention for LLMs' and 'Implementing Causal Self-Attention in PyTorch'. Pay close attention to how the masking is implemented with torch.tril and masked_fill before the softmax is applied. This directly translates the conceptual masking into efficient tensor operations.

Test your understanding!

In the previous lesson on BERT, we learned that it uses bidirectional self-attention. In this lesson, we learned GPT uses causal self-attention. Imagine you replaced the causal attention in a GPT model with BERT's bidirectional attention. What would happen during the pre-training task (next-token prediction)? Would the model learn to generate language effectively?

Show answer

If you used bidirectional attention, the model would perform the next-token prediction task perfectly, but it would not learn anything useful about generating language.

Reasoning: During training, to predict the token at position t, a model with bidirectional attention could simply "look ahead" to the token at position t in the input and copy it. Its attention mechanism would learn to assign a weight of 1.0 to the token it's supposed to be predicting. The training loss would quickly drop to zero, but the model wouldn't have learned the underlying patterns of language necessary to generate a new token when that future information isn't available (as is the case during actual generation/inference). Causal masking is essential because it forces the model to make its prediction based only on the information it has already seen.

3. Assembling the Full Architecture: Evolution from GPT-1 to GPT-3

A complete GPT model is constructed by stacking these decoder blocks on top of each other. Each block contains the core components we've discussed:

  1. Masked Multi-Head Self-Attention: For communication between tokens.
  2. Position-wise Feed-Forward Network: For computation on each token's representation.
  3. Residual Connections: To allow gradients to flow easily through the deep network.
  4. Layer Normalization: To stabilize the training process.

Interestingly, the evolution from GPT-1 to GPT-3 didn't involve radical architectural redesigns. The core blueprint remained remarkably consistent. The primary change was scale.

Let's look at the specifications of the GPT family to understand this progression.

karpathy/minGPT: A minimal PyTorch re-implementation ...

The minGPT repository by Andrej Karpathy contains excellent summaries of the key architectural parameters for the GPT family, extracted from the original papers. This will give us concrete numbers to analyze.

Read the three short sections detailing the architectures for GPT-1, GPT-2, and GPT-3. Pay attention to how the number of layers, model dimension (d_model), number of heads, and context size change across versions.

Here's a summary of that evolution:

Parameter GPT-1 (2018) GPT-2 (2019) GPT-3 (2020)
Parameters 117 Million 1.5 Billion 175 Billion
Layers 12 48 96
Model Dim (d_model) 768 1600 12,288
Heads 12 (not specified, but large) 96
Context Size 512 tokens 1024 tokens 2048 tokens
Key Arch. Tweak - LayerNorm moved to sub-block input (pre-norm) Architecture same as GPT-2

The most notable architectural tweak was in GPT-2, which moved Layer Normalization to the input of each sub-block (a "pre-norm" configuration), a practice that has since become standard as it tends to improve training stability. Otherwise, the story of GPT's development is a story of scaling up every parameter imaginable.

Conclusion

In this lesson, we've analyzed the architecture of GPT, the workhorse of modern generative AI. We contrasted its decoder-only design with BERT's encoder-only structure, establishing its specialization for autoregressive text generation.

  • Key Takeaways:
    • GPT models are decoder-only Transformers designed for autoregressive next-token prediction.
    • The core mechanism enabling this is causal self-attention, which uses masking to prevent a token from attending to future tokens in the sequence.
    • This masking is implemented efficiently by applying a lower-triangular mask to the attention score matrix before the softmax operation.
    • A full GPT model consists of input embeddings (token + position), a stack of decoder blocks (causal attention + FFN), and a final linear layer to map the output to vocabulary probabilities.
    • The evolution from GPT-1 to GPT-3 was primarily driven by a massive increase in scale (parameters, data, context window) rather than fundamental architectural changes.

Preview of the Next Lesson:
We've seen that the key to GPT's power lies in its scale. This wasn't just a brute-force effort; it was guided by a surprisingly predictable science. In our next lesson, we will delve into the "Scaling Laws for Large Language Models", exploring the empirical research that revealed the mathematical relationship between model size, dataset size, and performance, which provided the roadmap for building models like GPT-3.

Can't find a good explanation? Sign up and we'll make it for you

Sign up