Hello! Welcome to the next lesson in our journey through the Transformer architecture.
In our previous lesson, we built the multi-head attention layer. This powerful mechanism allows a model to weigh the importance of different words in a sequence when producing a representation for a specific word. We saw how using multiple "heads" in parallel enables the model to capture diverse syntactic and semantic relationships simultaneously.
However, the self-attention mechanism we built has a fundamental limitation: it is permutation-invariant. It treats the input sequence as a "bag of tokens" without any inherent sense of order. For example, to a pure self-attention layer, the sentences "The dog chased the cat" and "The cat chased the dog" are almost indistinguishable, as the pairwise attention scores between "dog" and "cat" would be similar in both cases.
This insensitivity to order is a major problem for tasks where sequence is critical, like language understanding. Today, we will solve this problem. Your learning outcome is to implement sinusoidal positional encodings for sequence order information. We will inject explicit information about the position of each token into the model, giving it the sequential context it needs.
1. The Problem: Permutation Invariance
Before we dive into the solution, let's solidify our understanding of the problem. As you know from your computer science background, a set is an unordered collection. The self-attention mechanism, by its design, operates on the input sequence more like a set than an ordered list.
To understand why, let's recall the core calculation: . The score between any two tokens depends only on the dot product of their query and key vectors, not their absolute or relative positions. If we shuffled the input embeddings, the resulting set of output vectors from the attention layer would be the same, just in a shuffled order.
The following resource provides a concise explanation of this challenge.
Positional Encodings Explained
This article from CodeSignal Learn clearly articulates why sequence order is lost in the self-attention mechanism.
Read the section titled 'The Permutation Invariance Problem'. It explains exactly why a mechanism that relies on query-key similarity is blind to the order of the inputs.
To overcome this, the authors of the original Transformer paper introduced a simple but brilliant idea: Positional Encodings.
2. The Solution: Sinusoidal Positional Encodings
The idea is to create a vector, the "positional encoding," that contains information about a token's position in the sequence. This vector has the same dimension as the token embedding (d_model), and we simply add it to the token embedding before feeding the sequence into the Transformer's encoder or decoder stack.

There are many ways to create these encodings. One could even have the model learn them using an nn.Embedding layer. However, the original paper proposed a fixed method using sine and cosine functions of different frequencies, which has proven to be very effective.
The formulas are as follows:
Let's break this down:
pos: The position of the token in the sequence (e.g., 0, 1, 2, ...).i: The index for the dimension pairs. It goes from0tod_model/2 - 1.2iand2i+1: These refer to the even and odd indices of the embedding dimension, respectively. So, for each pair of dimensions, we compute a sine and a cosine value.d_model: The dimension of the model's embeddings (e.g., 512).
The core of the formula is the term inside the sine/cosine functions. The denominator, , creates a geometric progression of wavelengths, from to . This means:
- For small values of
i(the first dimensions of the embedding), the wavelength is short. The positional encoding values change rapidly withpos. - For large values of
i(the last dimensions of the embedding), the wavelength is long. The positional encoding values change very slowly withpos.
This design has several elegant properties:
- Unique Encoding: Each position
posgets a unique encoding vector. - Extrapolation: The sinusoidal nature allows the model to potentially understand the order of sequences longer than any it saw during training.
- Relative Positions: A key benefit is that for any fixed offset
k, can be represented as a linear function of . This makes it easy for the model to learn to attend based on relative positions (e.g., "the word 3 positions to my left").
The following video provides an excellent conceptual and mathematical overview.
Positional Encoding in Transformer Neural Networks Explained
This video from CodeEmporium explains the motivation and the mathematical formulation behind sinusoidal positional encodings.
Watch the segment from 05:03 to 07:26. It clearly explains the formula and the three key reasons for this specific design: periodicity, constrained values, and extrapolation.
3. Implementation from Scratch
Now for the main objective: implementing this in PyTorch. Our goal is to create a PositionalEncoding module that can be integrated into our Transformer model. The implementation will follow the formulas we just discussed.
The most common and numerically stable way to implement this is to compute the denominator term in log-space, which is equivalent to the original formula:div_term =
This video provides a fantastic, step-by-step walkthrough of how to build the positional encoding matrix from the ground up in PyTorch.
Positional Encoding in Transformer Neural Networks Explained
Let's return to the CodeEmporium video. The next section is a complete, line-by-line coding tutorial that perfectly matches our learning objective.
Watch the implementation section from 07:26 to 11:44. The presenter builds the logic step-by-step before wrapping it in a reusable PyTorch nn.Module. Pay close attention to: The calculation of the denominator (our div_term). The creation of the position tensor. How torch.sin and torch.cos are applied to the even and odd dimensions using slicing (pe[:, 0::2], pe[:, 1::2]). The final interleaving/flattening to get the final matrix.
The video builds the logic interactively. The final, polished implementation is often encapsulated in a class, as shown in the "Annotated Transformer," a famous guide that pairs the original paper with PyTorch code.
The 'Annotated Transformer' provides the canonical PyTorch implementation of the PositionalEncoding module. This is what you would typically see in production codebases.
Read the section '3.4 Positional Encoding'. Focus on the PositionalEncoding class implementation. Notice the use of self.register_buffer('pe', pe). This is a key PyTorch feature that makes pe part of the module's state (so it's saved with the model and moved to GPU with .to(device)), but not a parameter that gets updated by the optimizer.
The forward pass of the module is simple but crucial: it takes the input embeddings x and adds the pre-computed positional encodings. A dropout layer is also typically applied.
def forward(self, x):
# x has shape (batch_size, seq_len, d_model)
# self.pe has shape (1, max_len, d_model)
# Add the positional encoding to the input embeddings
# We slice self.pe to match the sequence length of the input x
x = x + self.pe[:, :x.size(1)]
# Apply dropout for regularization
return self.dropout(x)
Test your understanding!
Consider the positional encoding formula. You have an embedding of size d_model = 512.
- For the token at position
pos = 20, which function (sine or cosine) would be used to calculate the value for the 100th dimension (i.e., at index 99 of the vector)? - Which function would be used for the 101st dimension (at index 100)?
- Why is it important to use
register_bufferinstead of just assigningself.pe = pein the__init__method of a PyTorchnn.Module?
Show answer
- The 100th dimension corresponds to vector index 99. Since 99 is odd, it falls under the
2i+1case. Therefore, cosine would be used. - The 101st dimension corresponds to vector index 100. Since 100 is even, it falls under the
2icase. Therefore, sine would be used. register_bufferensures that thepetensor is considered part of the module's state. This means that when you call.to(device)on your model, the buffer will also be moved to the correct device (e.g., GPU). It also ensures the tensor is saved in the model'sstate_dict, but it is not registered as a parameter, so it won't be updated during training by the optimizer. A simple Python attribute assignment (self.pe = pe) would not have these benefits.
Conclusion
In this lesson, we tackled a critical limitation of the self-attention mechanism and implemented the elegant solution proposed by the Transformer architecture. By injecting fixed, sinusoidal positional information directly into the input embeddings, we give the model the awareness of sequence order it needs to understand complex dependencies in data.
Key Takeaways:
- Self-attention and multi-head attention are permutation-invariant, meaning they ignore the order of tokens in a sequence.
- Sinusoidal Positional Encodings are fixed vectors added to token embeddings to provide this missing order information.
- The encoding for each position is a unique vector generated by sine and cosine functions of varying frequencies across the embedding dimensions.
- The implementation involves pre-computing a matrix of these encodings up to a maximum length and storing it as a non-trainable buffer in a PyTorch module.
Preview of the Next Lesson:
We are now equipped with all the necessary sub-components! We have:
- An embedding layer to convert tokens to vectors.
- A positional encoding layer to add sequence order information.
- A multi-head attention layer to model relationships between tokens.
In the next lesson, "Build a complete Transformer encoder block," we will assemble these pieces. We will combine multi-head attention with a position-wise feed-forward network, using residual connections and layer normalization to tie everything together and ensure stable training. This will give us the complete, repeatable block that forms the backbone of the Transformer encoder.