Hello! Welcome to the first lesson of our module on model optimization.
In our last lesson, we concluded our tour of advanced audio language models by looking at the significant challenges that define the research frontier. We saw that while models like AudioPaLM are incredibly powerful, their enormous size and computational cost create major barriers to deployment and research. This provides the perfect motivation for our current focus: making these powerful models smaller, faster, and more efficient without sacrificing too much performance.
Today, we will dive into one of the most effective techniques for this: knowledge distillation. This lesson will explain the core principles behind this "teacher-student" training paradigm. You will learn how a large, complex model (the "teacher") can transfer its "knowledge" to a smaller, faster "student" model. We will then see a direct application of this by exploring how the popular Distil-Whisper model was created by distilling the larger Whisper model.
1. The Motivation: Why Distill Knowledge?
The core challenge in deploying large neural networks is a fundamental conflict between the needs of training and inference.
- During training, our main goal is to maximize accuracy. We achieve this by using large models with many parameters, training on vast datasets, and using techniques like ensembling. This results in powerful but "cumbersome" models that are slow and resource-intensive.
- During inference (deployment), our priorities shift. We need models that are not only accurate but also fast, have low latency, and can run on devices with limited computational power and memory (like a smartphone or an edge device).
Knowledge distillation aims to resolve this conflict by transferring the knowledge from a large, trained model to a smaller one that is better suited for deployment.
To introduce this idea, let's watch a short video that uses a great analogy.
Distilling the Knowledge in a Neural Network
This clip from the video "Distilling the Knowledge in a Neural Network" by Kapil Sachdeva introduces the core motivation for knowledge distillation, comparing the different requirements of training and inference to the life stages of an organism.
Watch from 00:00:35 to 00:03:12. Pay attention to the distinction between a 'cumbersome model' (the teacher) and a 'simpler model' (the student).
This concept, formally introduced by Geoffrey Hinton and his colleagues, has become a cornerstone of model compression. As we'll see, it's particularly effective in the audio domain. For instance, the Distil-Whisper model achieves a nearly 6x speed-up over the original Whisper-large model with only a minor drop in accuracy, making it far more practical for real-world applications.
2. The Core Principles of Knowledge Distillation
How can we make a small student model learn from a large teacher? The key lies in changing what the student model learns.
2.1 Beyond Hard Labels: The Value of "Soft Targets"
In standard classification training, a model learns from hard labels. For example, in a speech recognition task, the ground truth is a sequence of tokens. The loss function (typically cross-entropy) penalizes the model for not predicting the correct tokens. This approach tells the model what the single correct answer is, but it provides no information about the relationships between incorrect answers.
For example, when a teacher model hears the word "cat," it might assign a high probability to "cat" (e.g., 80%), but it might also assign a small but non-zero probability to a similar-sounding word like "cap" (e.g., 5%) and a near-zero probability to a dissimilar word like "table." This nuanced distribution is what the original paper calls "dark knowledge." Hard labels discard this rich information.
Knowledge distillation leverages this by training the student on the teacher's full output probability distribution, known as soft targets.
2.2 The Role of Temperature
To better expose this "dark knowledge," we use a modified softmax function with a temperature parameter, . The standard softmax function is:
where are the logits (raw scores) from the model's final layer.
The temperature-scaled softmax is:
- When , we have the standard softmax.
- When , the distribution becomes "softer" or smoother. The probabilities are less concentrated on the highest-scoring logit, revealing the relative similarities the teacher model has learned between classes.
The following video provides an excellent mathematical and visual intuition for the softmax function and the effect of temperature.
Distilling the Knowledge in a Neural Network
Let's watch another segment from Kapil Sachdeva's video. This part details the softmax function and, crucially, explains how the temperature parameter helps create the 'soft labels' needed for distillation.
Watch from 00:05:22 to 00:14:03. Focus on: Why the softmax function is used to create a probability distribution. How a high temperature T 'smoothes' the distribution, revealing relative similarities between classes that a standard softmax would hide.
2.3 The Distillation Loss Function
The student model is trained to minimize the difference between its own output distribution and the teacher's soft targets. This is typically done using the Kullback-Leibler (KL) Divergence, a measure of how one probability distribution differs from a second, reference distribution.
The overall training process often uses a composite loss function that is a weighted sum of two components:
- Distillation Loss: A KL Divergence loss between the student's soft predictions (calculated with ) and the teacher's soft targets (also calculated with ). This encourages the student to mimic the teacher's reasoning.
- Student Loss: A standard cross-entropy loss between the student's hard predictions (calculated with ) and the ground-truth hard labels. This anchors the student to the correct answers.
The final loss can be expressed as:
where is a hyperparameter balancing the two terms. The term is a scaling factor used in the original paper to ensure the gradient magnitudes from the soft and hard targets are roughly on the same scale.
This entire process is beautifully illustrated in the following diagram.

3. Application: How to Distill a Speech Model
Now let's move from theory to practice. Given your background with Whisper, the Distil-Whisper project is a perfect case study. It follows a clear, multi-stage process that you can adapt for your own projects.
The official Hugging Face repository provides scripts that make this process accessible. We'll walk through the three main stages conceptually.
supawichwac/training - Hugging Face
The supawichwac/training repository on Hugging Face contains the PyTorch scripts for reproducing Distil-Whisper. We will use its README as a guide for our practical application walkthrough.
First, read the introduction section of the README. This will confirm that the scripts are in PyTorch and designed for distilling Whisper on custom languages/datasets.
Stage 1: Pseudo-Labelling (Optional but Powerful)
To achieve the best results, distillation requires a large amount of training data. A powerful strategy is to use a massive corpus of unlabeled audio. The large teacher model (e.g., openai/whisper-large-v3) is used to transcribe this audio, creating a vast dataset of (audio, pseudo-label) pairs. The student is then trained on this dataset. This allows the student to learn from a much wider variety of audio than might be available in a curated, human-labeled dataset.
supawichwac/training - Hugging Face
Let's examine the first step in the Distil-Whisper pipeline: generating pseudo-labels with the teacher model.
Read 'Section 1. Pseudo-Labelling'. Note how the run_pseudo_labelling.py script takes a teacher model and a dataset to generate transcriptions. This step effectively scales up the training data using the teacher's knowledge.
Stage 2: Student Model Initialization
We don't start with a randomly initialized student model. Instead, we create a smaller version of the teacher by pruning, or removing, some of its layers. A common and effective strategy is to copy a subset of the teacher's layers, spaced as far apart as possible, into the student. For Distil-Whisper, the student retains all 32 encoder layers of Whisper-large-v3 but only keeps 2 of the 32 decoder layers. This preserves the powerful audio encoding capabilities of the teacher while drastically reducing the size and complexity of the autoregressive decoder.
supawichwac/training - Hugging Face
The next step is to create the student model architecture and initialize its weights from the teacher.
Read 'Section 2. Initialisation'. Pay attention to how create_student_model.py works: it copies 'maximally spaced layers' from the teacher to initialize the student. This is a much better starting point than random initialization.
Stage 3: Distillation Training
This is where the core learning happens. The initialized student model is trained on the pseudo-labeled dataset using the composite loss function we discussed earlier. In each training step:
- A batch of audio is fed to both the teacher and the student.
- The teacher, running in evaluation mode (
no_grad()), generates the soft targets. - The student generates its own predictions.
- The distillation loss (KL Divergence) and student loss (Cross-Entropy) are calculated and combined.
- The gradients are computed, and the student model's weights are updated.
The following resources explain this training stage and the associated script.
supawichwac/training - Hugging Face
Finally, we perform the actual distillation training.
Read 'Section 3. Training'. Notice how run_distillation.py brings everything together: it takes the student model from Stage 2, the teacher model, and the pseudo-labeled data from Stage 1 to perform the final training.
For a more in-depth look at a distillation script, the following video provides a code walkthrough. While the example is for an LLM, the principles—loading a teacher and student, setting the teacher to eval mode, and implementing a custom loss function—are directly transferable to a speech model like Whisper.
Distillation of Transformer Models
Let's watch a code walkthrough for a distillation script. This will help solidify your understanding of how the theoretical concepts are implemented in a PyTorch training loop.
Watch the segment from 00:39:23 to 00:47:59. Focus on these key implementation details: Loading both a teacher and a student model (00:44:42). Setting the teacher to evaluation mode to prevent its weights from being updated (00:44:56). Implementing a custom compute_loss function that calculates the KL Divergence between teacher and student logits (00:45:50). The use of a temperature parameter within the loss function (00:47:30).
4. Types of Knowledge Distillation
What we've just covered is the classic and most common form of knowledge distillation, often called response-based distillation, because the student learns from the teacher's final output (its response).
However, for a research-oriented perspective, it's useful to know that there are other forms of knowledge that can be transferred.
Everything You Need to Know about Knowledge Distillation
This blog post from Hugging Face provides an excellent taxonomy of different distillation methods.
Read the sections 'Types of knowledge distillation' and 'Improved algorithms'. This will introduce you to concepts like: Feature-based distillation: The student mimics the teacher's intermediate layer activations. Relation-based distillation: The student learns the relationships between data points, as seen by the teacher. Self-distillation: A model learns from itself, where a deeper layer acts as the teacher for a shallower layer. Quantized distillation: A method that combines distillation with quantization, another optimization technique we will cover next.
Understanding these variations is valuable as they offer different tools to tackle specific compression challenges and are active areas of research.
Conclusion
In this lesson, we have demystified the process of knowledge distillation. You now understand not only the theoretical principles but also how they are applied in practice to create smaller, more efficient speech models like Distil-Whisper.
Key Takeaways:
- Motivation: Distillation bridges the gap between large, accurate models needed for training and small, fast models required for deployment.
- Core Principle: Instead of learning from hard labels alone, the student model learns from the teacher's rich, nuanced probability distributions (soft targets).
- Key Components: The process relies on a temperature-scaled softmax to create soft targets and a composite loss function, typically combining KL Divergence (for mimicking the teacher) and Cross-Entropy (for matching the ground truth).
- Application to Speech: A common pipeline involves (1) using the teacher to pseudo-label a large unlabeled dataset, (2) initializing a smaller student by pruning the teacher, and (3) performing distillation training.
Preview of the Next Lesson:
Knowledge distillation is a powerful training-time optimization technique. In our next lesson, we will explore quantization, a popular post-training optimization method. You'll learn how to further reduce a model's size and speed up inference by representing its weights with lower-precision numbers. This technique can even be applied to a model that has already been distilled, offering another layer of optimization.