Hello! Welcome to the next lesson in our module, "Foundations of Language Modeling and Embeddings."
In our previous lesson, we built a Word2Vec model to learn dense vector representations for words. This was a huge step up from one-hot encoding, as it captured semantic meaning. However, the model operated on a fixed vocabulary of whole words. This leads to a critical problem: how do we handle words that weren't in our training data (Out-Of-Vocabulary or OOV words)? Furthermore, how can we represent the relationship between words like "run", "running", and "ran" in a more systematic way? And what about languages like your N4-level Japanese, where words aren't separated by spaces?
Today, we will address these challenges. The learning outcome for this lesson is to apply tokenization techniques, including Byte-Pair Encoding (BPE) and WordPiece. We'll move beyond word-level processing and explore the powerful world of subword tokenization, which is the standard for virtually all modern language models.
By the end of this lesson, you will understand:
- The motivation for subword tokenization over word-based or character-based approaches.
- The step-by-step algorithm behind Byte-Pair Encoding (BPE).
- How WordPiece is similar to BPE, and the crucial difference in its merge criterion.
- How to train and use a custom tokenizer using the popular Hugging Face
tokenizerslibrary.
1. From Words to Subwords: The Motivation
The core idea of tokenization is to break a piece of text into smaller units called tokens. These tokens are then mapped to integer IDs, which are subsequently fed into an embedding layer, just like we saw with Word2Vec.
- Word Tokenization: Simple, but fails on any word not in its vocabulary. It also treats related words like "help" and "helpful" as completely separate, missing their shared root.
- Character Tokenization: Solves the OOV problem, as any word can be built from characters. However, it creates very long sequences, losing the semantic meaning of whole words and increasing the computational burden on the model.
Subword tokenization is the effective middle ground. It breaks text down into chunks that are often smaller than words but larger than single characters. Common words can remain as single tokens, while rare words are broken down into smaller, meaningful subword units. For example, "tokenization" might become ["token", "ization"]. This allows the model to handle unseen words like "hypertokenization" by breaking them into familiar pieces: ["hyper", "token", "ization"].
LLM Tokenizers Explained: BPE Encoding, WordPiece and SentencePiece
To start, let's get a quick overview of what tokenization is and the main algorithms we'll be discussing today.
Watch the first 33 seconds of the video "LLM Tokenizers Explained" by DataMListic. It provides a concise summary of the role of a tokenizer in the language model pipeline.
2. Byte-Pair Encoding (BPE)
Byte-Pair Encoding was originally a data compression algorithm. Its adaptation to NLP, popularized by the GPT-2 paper, was a major breakthrough. The algorithm is elegant and data-driven.
The process for training a BPE tokenizer is as follows:
- Initialize Vocabulary: Start with a vocabulary containing all the individual characters (or bytes) present in your training corpus.
- Iterative Merging:
a. Count the frequency of all adjacent pairs of tokens in the corpus.
b. Find the most frequent pair (e.g.,'e', 'r').
c. Merge this pair into a single new token (e.g.,'er').
d. Add this new token to your vocabulary and the merge rule to a merge list. - Repeat: Continue this process for a predetermined number of merges. This number, combined with the initial character set, determines your final vocabulary size.
Let's walk through a simple example. Suppose our corpus is: new, newest, wide, widest.
- Initial vocab:
n, e, w, s, t, i, d - Step 1: The most frequent pair is
'e', 's'(appears twice innewestandwidest). We merge them to create'es'.- Our vocabulary is now
n, e, w, s, t, i, d, es. - Our text is effectively
new, new(es)t, wide, wid(es)t.
- Our vocabulary is now
- Step 2: The next most frequent pair is
'es', 't'(appears twice). We merge them to create'est'.- Our vocabulary is now
n, e, w, s, t, i, d, es, est. - Our text is now
new, new(est), wide, wid(est).
- Our vocabulary is now
- And so on.
LLM Tokenizers Explained: BPE Encoding, WordPiece and SentencePiece
The same video provides an excellent animated explanation of this BPE process. Watching it will solidify your understanding of the iterative merging.
Watch from 00:33 to 02:03. Notice how the vocabulary is built up step-by-step by merging the most frequent consecutive characters.
For a code-centric perspective, Andrej Karpathy's minbpe is a fantastic resource. It's a minimal, clean implementation of BPE.
karpathy/minbpe: Minimal, clean code for the Byte Pair ...
Let's look at a practical, coded example from the minbpe repository. This will show you how the BPE algorithm translates directly into a simple Python script.
Read the 'quick start' section. It demonstrates training a basic tokenizer and encoding text, reproducing the Wikipedia example of BPE. Pay attention to how the token IDs are assigned, starting from 256 (after the initial 256 byte values).
3. WordPiece: A Likelihood-Based Approach
The WordPiece algorithm, used by influential models like BERT, is very similar to BPE but with one critical difference in the merging strategy.
- BPE: Merges the most frequent pair of tokens.
- WordPiece: Merges the pair that maximizes the likelihood of the language model once added to the vocabulary.
What does "maximizing likelihood" mean in this context? A simplified way to think about it is that WordPiece chooses to merge a pair of tokens (say, A and B) if the probability of the merged token AB is significantly higher than the product of the probabilities of the individual tokens A and B. This is calculated from their counts in the training data: score = count(AB) / (count(A) * count(B)). This score favors pairs that co-occur frequently but are not just two individually frequent tokens happening to be next to each other.

Another key characteristic of WordPiece is its notation. When a word is split, the subsequent pieces are prefixed with ## to indicate they are continuations of a word.
tokenization->["token", "##ization"]
This prefixing is important for unambiguously reconstructing the original text.
LLM Tokenizers Explained: BPE Encoding, WordPiece and SentencePiece
Let's return to the DataMListic video for a concise explanation of how WordPiece's merging criterion differs from BPE's.
Watch from 02:03 to 02:50. Focus on the core difference: WordPiece doesn't choose the most frequent pair but the one that 'maximizes the likelihood of the training data'.
4. Application: Training a WordPiece Tokenizer
Given your background as a software engineer, the best way to solidify these concepts is to see them in action. We'll look at how to train a tokenizer from scratch using Hugging Face's tokenizers library, which is the industry standard.
This library provides highly optimized implementations of BPE, WordPiece, and other tokenization algorithms. We'll focus on BertWordPieceTokenizer.
How to Build a Bert WordPiece Tokenizer in Python and HuggingFace
The following video by James Briggs is a practical, step-by-step guide to building a WordPiece tokenizer. We will focus on the tokenizer training and usage parts.
Please watch the following two segments: Training the Tokenizer (17:38 - 24:51): Pay close attention to the initialization of BertWordPieceTokenizer and the parameters passed to the .train() method, such as vocab_size, min_frequency, and special_tokens. Loading and Using the Tokenizer (24:51 - 31:10): Observe how the trained tokenizer is loaded and then used to tokenize sentences. Note how it handles OOV words by breaking them down into known word pieces, as demonstrated with the word responsabilità .
The process shown in the video is a standard workflow in NLP:
- Gather a Corpus: Collect a large amount of raw text.
- Initialize a Tokenizer: Choose an algorithm (e.g., WordPiece) and set its parameters.
- Train: Feed the corpus to the tokenizer's
trainmethod. It will learn the merge rules and build the vocabulary. - Save & Use: Save the vocabulary and merge rules to a file. You can then load this tokenizer to consistently convert new text into token IDs for your model.
Test your understanding!
You are training a new tokenizer. Your training data contains the word "unhappiness" many times.
- Using a BPE or WordPiece tokenizer, what is a likely way for "unhappiness" to be tokenized after training?
- Why is this subword tokenization more useful for the model than treating "unhappiness" as a single, unique token?
- You encounter the word "unhappily" in your test data, which never appeared during training. How would your trained tokenizer likely handle it?
Show answer
- A likely tokenization would be
["un", "##happy", "##ness"]. The tokenizer would learn that "un", "happy", and "ness" are common and meaningful morphemes. - By breaking it down, the model can learn a general representation for the prefix "un-" (negation), the root "happy" (emotion), and the suffix "-ness" (state of being). This knowledge is transferable. It can help the model understand other words like "unhelpful" or "sadness" even if it has seen them less frequently.
- The tokenizer would likely handle it gracefully by tokenizing it as
["un", "##happy", "##ly"]. Since it already has tokens for "un" and "happy", it only needs to have seen the suffix "ly" in other contexts (e.g., in "quickly", "softly") to be able to represent this new word. This demonstrates the power of subword tokenization in handling OOV words.
5. SentencePiece and Other Considerations
A challenge with both BPE and WordPiece as originally defined is that they rely on some form of pre-tokenization (e.g., splitting by spaces and punctuation). This can be problematic for languages that don't use spaces, like Japanese or Chinese.
SentencePiece is a technique that solves this by treating the input text as a single stream of Unicode characters. It encodes whitespace directly into the tokens, often using a special character like (U+2581) to represent a space. This makes the tokenization and detokenization process perfectly reversible without relying on language-specific rules.
LLM Tokenizers Explained: BPE Encoding, WordPiece and SentencePiece
The final segment of the DataMListic video introduces SentencePiece and its advantages, which should be particularly relevant given your interest in Japanese.
Watch from 03:16 to 04:25. Focus on how SentencePiece treats the input as a raw stream of characters, including whitespace, which solves pre-tokenization issues for languages like Japanese and Chinese.
Conclusion
Today we've dismantled the black box of tokenization, a fundamental process for all modern language models. You've seen how we've evolved from fragile word-based vocabularies to robust and efficient subword systems.
Key Takeaways:
- Subword tokenization provides a balance between word-level semantics and character-level flexibility, solving the Out-Of-Vocabulary (OOV) problem.
- Byte-Pair Encoding (BPE) is a data-driven algorithm that builds a vocabulary by iteratively merging the most frequent adjacent pairs of tokens.
- WordPiece is similar to BPE but uses a probabilistic criterion, merging pairs that maximize the likelihood of the training data.
- Practical tokenization is done with highly optimized libraries like Hugging Face's
tokenizers, which allow you to train custom tokenizers on your own data. - SentencePiece enhances these algorithms by removing the need for language-specific pre-tokenization, making it more versatile across different languages.
Preview of the Next Lesson:
Now that we know how to convert raw text into a sequence of subword token IDs, we are ready to explore how models like BERT learn from this input. In the next lesson, we will delve into the masked language modeling (MLM) objective, the clever self-supervised training task that allows BERT to learn deep contextual representations of language. You'll see exactly how the WordPiece tokens we've discussed today are used in this process.