Hello! Welcome to the next lesson in our journey through the foundations of language modeling.
In our last session, we explored how raw text is broken down into subword tokens using techniques like WordPiece. This was a crucial first step, as it provides a robust way to convert any text into a sequence of numerical IDs for a model. Now, we'll address the next logical question: how does a model like BERT actually learn from these token sequences to understand language?
Today's learning outcome is to explain the masked language modeling (MLM) objective used by BERT. This is the self-supervised training task that was central to BERT's success and distinguished it from previous language models.
By the end of this lesson, you will understand:
- The core idea behind BERT's bidirectionality and how it differs from left-to-right models.
- The mechanics of the Masked Language Modeling (MLM) objective.
- The clever "80-10-10" masking strategy and the rationale behind it.
- How the model uses the output of masked positions to calculate loss and learn.
- The role of the second pre-training task, Next Sentence Prediction (NSP), in the original BERT model.
1. The Core Idea: Bidirectionality
Before BERT, dominant language models like those based on LSTMs or the early GPT were causal or autoregressive. They processed text sequentially from left to right, predicting the next word based only on the words that came before it. This is like reading a sentence through a narrow slit, only seeing the past.
BERT (Bidirectional Encoder Representations from Transformers) introduced a paradigm shift. It's based on the encoder part of the Transformer architecture and is designed to be bidirectional, meaning it considers the entire sentence—both left and right context—at once to understand a word.
Instead of predicting the next word, BERT is pre-trained on a "fill-in-the-blanks" task, formally known as a cloze task. This is the essence of Masked Language Modeling.
A Dive into BERT's Masked Language Modeling
To build your intuition, let's read a short article that explains this fundamental concept of bidirectionality very clearly.
Read the introduction and the section 'The Core Intuition: Bidirectionality'. This will give you a strong conceptual foundation for why looking at both the left and right context is so powerful.
This self-supervised approach—creating a learning problem directly from the unlabeled input text—allows the model to learn a deep understanding of language structure and semantics without needing manually labeled data.
2. The MLM Training Process in Detail
So, how do we create this "fill-in-the-blanks" task for the model? It's more sophisticated than just hiding a word. The process involves randomly selecting tokens and applying a specific set of rules.
Here's the procedure:
- Take an input sentence and tokenize it using an algorithm like WordPiece (as we discussed in the last lesson).
- Randomly select 15% of the tokens from the sequence for prediction.
- For each selected token, apply the following "80-10-10" rule:
- 80% of the time: Replace the token with the special
[MASK]token.- Example:
The chef cooked a delicious [MASK] for the guests.-> The model must predictmeal.
- Example:
- 10% of the time: Replace the token with a random token from the vocabulary.
- Example:
The chef cooked a delicious shoe for the guests.-> The model must still predictmeal. This forces the model to learn that the given sentence is nonsensical and to rely on the surrounding context to correct it.
- Example:
- 10% of the time: Keep the token as is.
- Example:
The chef cooked a delicious meal for the guests.-> The model must still predictmeal. This biases the model towards the correct representation of the word.
- Example:
- 80% of the time: Replace the token with the special
Why this complex strategy? If we only used the [MASK] token, the model would become very good at predicting words when it sees [MASK], but it would never encounter this token during real-world use (fine-tuning). This creates a mismatch. The 10% random and 10% unchanged portions force the model to learn a rich, contextual representation for every token, not just the [MASK] token.
BERT explained: Training, Inference, BERT vs GPT/LLamA, Fine tuning, [CLS] token
Let's watch a video that walks through this masking procedure and explains how BERT's architecture facilitates it.
Watch the segment from 37:06 to 40:52. Pay close attention to the explanation of the 80-10-10 rule and the key point about how BERT is bidirectional because it does not use a causal mask in its self-attention mechanism. This allows every token to attend to every other token.

3. Prediction and Learning
Once the input is prepared, it's passed through the stack of Transformer encoder layers in BERT. The model's task is to predict the original tokens at the 15% of positions that were manipulated.
Here's how the learning happens:
- For each of the 15% selected positions, the model's final-layer output vector (a rich, contextualized embedding) is taken.
- This vector is passed through a small feed-forward network and then a softmax layer. This produces a probability distribution over the entire vocabulary (e.g., 30,000 possible WordPiece tokens).
- The cross-entropy loss is calculated between the model's predicted probability distribution and the actual, original token.
- Crucially, the loss is only computed for the 15% of tokens that were selected for masking. The model's predictions for the other 85% of tokens are completely ignored in the loss calculation. This focuses the training signal entirely on the "fill-in-the-blanks" task.
- The total loss is the average loss over all the masked positions, and this is what's used to update the model's weights via backpropagation.
BERT Neural Network - EXPLAINED!
For a clear breakdown of the output and loss calculation stages, the 'BERT Neural Network - EXPLAINED!' video is excellent.
Watch the segment from 08:18 to 09:46. Focus on how the output word vectors are converted to probability distributions and how the loss is calculated only for the masked words.
Test your understanding!
Imagine you are designing the BERT pre-training strategy. A colleague suggests simplifying the MLM objective by always replacing the chosen 15% of tokens with [MASK].
Why would you argue against this simplification? What specific problem would it cause when you later try to use the pre-trained BERT model for a downstream task like sentiment analysis?
Show answer
The problem is the pre-train/fine-tune mismatch. The model would be pre-trained on inputs containing the [MASK] token, but this token does not appear in the natural text used for fine-tuning (e.g., sentences for sentiment classification).
This means the model would learn powerful contextual representations for the tokens surrounding [MASK], but its representations for the actual words in the sentence might not be as robust. By including the 10% random replacement and 10% identity cases, we force the model to learn to critically evaluate and produce a good contextual representation for every token in the input sequence, making it more powerful and generalizable for downstream tasks where no [MASK] tokens are present.
4. The Other Half: Next Sentence Prediction (NSP)
The original BERT model was pre-trained on two tasks simultaneously: MLM and Next Sentence Prediction (NSP). While MLM taught the model about word-level relationships within a sentence, NSP was designed to teach it about relationships between sentences.
The NSP task is a simple binary classification:
- The model is given two sentences, A and B.
- 50% of the time, B is the actual sentence that follows A in the training corpus (a positive example).
- 50% of the time, B is a random sentence from a different document (a negative example).
- The model's task is to predict whether B is the true next sentence or not.
To facilitate this, special tokens are used:
[CLS]: A token added to the beginning of every input sequence. Its final hidden state is used to make the NSP prediction.[SEP]: A token used to separate sentence A from sentence B.

While NSP was part of the original BERT and important for its performance on certain tasks (like question answering), subsequent research (e.g., in the RoBERTa model) found that MLM alone was sufficient or even better. Nevertheless, understanding NSP is key to understanding the full context of BERT's design.
Fine-Tuning and Masked Language Models
The textbook chapter by Jurafsky and Martin provides a formal description of the NSP task.
Read the section 'Next Sentence Prediction' (Section 11.2.3). It clearly explains the motivation for NSP and the mechanics of how it's implemented using the [CLS] and [SEP] tokens.
Conclusion
In this lesson, we've unpacked the ingenious training objective that made BERT a landmark model in NLP. By creating a "fill-in-the-blanks" game for itself, BERT learned to understand language with a deep, bidirectional awareness of context.
Key Takeaways:
- BERT's key innovation is its bidirectional nature, allowing it to use both left and right context simultaneously.
- The Masked Language Modeling (MLM) objective trains the model by corrupting a percentage of input tokens and tasking the model with predicting their original values.
- The 80-10-10 masking rule is a crucial detail that avoids a mismatch between pre-training and fine-tuning, forcing the model to learn robust representations for all tokens.
- Learning is driven by calculating cross-entropy loss only on the positions that were masked or altered.
- The original BERT also used Next Sentence Prediction (NSP) to learn sentence-level relationships, though this task is less common in modern encoder models.
Preview of the Next Lesson:
We've now seen BERT's "corrupt and reconstruct" approach, which is ideal for creating rich representations for language understanding tasks. But what if our goal is language generation? For that, a different approach is needed. In our next lesson, we will explore the causal (autoregressive) language modeling objective used by GPT, which forms the foundation of modern generative LLMs. This will highlight the fundamental fork in the road between encoder-style models like BERT and decoder-style models like GPT.