Skip to main content
Create your own
Lesson illustration

Understanding HuBERT: Architecture and Pretraining

Hello! Welcome to the fourth lesson in our module on "Self-Supervised Speech Representation."

In our last two lessons, we delved into the wav2vec 2.0 framework. We saw how it uses a CNN and Transformer to create speech representations and, crucially, how its pre-training objective works. We learned that wav2vec 2.0 performs online discretization using a quantization module with the Gumbel-Softmax trick and learns via a complex contrastive loss function. This approach, while powerful, involves several moving parts and sensitive hyperparameters like the Gumbel-Softmax temperature.

Today, we'll explore HuBERT (Hidden-Unit BERT), a model that builds on the same architectural foundation as wav2vec 2.0 but radically simplifies the pre-training objective. HuBERT's innovation lies in decoupling the process of creating discrete targets from the representation learning task itself. This lesson directly addresses the learning outcome: Describe the HuBERT architecture and its offline clustering-based pretraining objective (masked prediction of acoustic units).

By the end of this lesson, you will understand:

  • The core architectural components of HuBERT.
  • The two-step pre-training process involving offline clustering and masked prediction.
  • How HuBERT iteratively refines its own targets to improve its representations.
  • The key philosophical and practical differences between HuBERT and wav2vec 2.0.

1. The Motivation: Simplifying Self-Supervised Learning for Speech

While wav2vec 2.0 demonstrated state-of-the-art results, its training objective is intricate. It requires simultaneously learning the feature representations and the discrete targets, balanced by a contrastive loss and a diversity loss. This joint learning can be difficult to tune.

The creators of HuBERT asked a key question: what if we separate these two problems?

  1. First, generate a set of discrete "acoustic unit" labels for the unlabeled audio.
  2. Then, train a standard BERT-like model to predict these labels in a masked prediction task.

This approach simplifies the objective to a standard cross-entropy loss, making the training process more stable and analogous to how BERT is trained on text. The main challenge shifts from designing a complex loss function to generating good-quality discrete targets from continuous speech.

To begin, let's watch a short video that introduces the core idea of HuBERT and how it relates to BERT.

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units #nlp

This video from the Social Robotics Talk channel provides an excellent introduction to HuBERT, breaking down the name and connecting its self-supervised method to the masked language modeling used in BERT.

Watch from 00:52 to 06:02. Pay attention to how the presenter explains self-supervision in the context of masked modeling and why speech presents a unique challenge (continuous input) that HuBERT aims to solve.


2. The HuBERT Architecture and Pre-training Framework

At a high level, the HuBERT architecture is nearly identical to wav2vec 2.0:

  • A multi-layer Convolutional Neural Network (CNN) acts as a feature encoder, processing the raw audio waveform and downsampling it into a sequence of feature vectors (typically at a 50Hz or 20ms frame rate).
  • A multi-layer Transformer (BERT) Encoder then takes this sequence of features and produces contextualized representations.

The innovation is in the training loop, which is an iterative, two-step process.

HuBERT Training Process Explained
This diagram from the blog 'HuBERT: How to Apply BERT to Speech, Visually Explained' illustrates the two-step training process. In Step 1, discrete 'hidden unit' targets are discovered via clustering. In Step 2, the main model learns to predict these targets for masked portions of the audio.

Let's dive into these two steps.

Step 1: Discovering Hidden Units via Offline Clustering

The first, crucial step is to generate discrete target labels from the continuous audio. HuBERT does this "offline" using a simple but effective clustering algorithm: K-means.

For the very first iteration of training:

  1. Feature Extraction: For each audio file in the training set, we first extract a standard acoustic feature, Mel-Frequency Cepstral Coefficients (MFCCs).
  2. Clustering: A K-means algorithm is trained on all these MFCC vectors. Let's say we use K=100 clusters.
  3. Label Generation: Each 20ms frame of audio, represented by its MFCC vector, is then assigned the ID of the closest cluster centroid (e.g., a label from 0 to 99). These cluster IDs are our initial "hidden units".

These hidden units are analogous to words in a text vocabulary. They are imperfect, "noisy" labels, as they are derived without any human annotation, but they are consistent. For example, similar-sounding vowels across the dataset will likely be grouped into the same cluster.

This next video segment walks through this first step in detail.

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units #nlp

Let's return to the Social Robotics Talk video, which now explains this first step of generating hidden units.

Watch from 07:33 to 12:59. This covers the necessity of the two-step process for speech and then details Step 1: using K-means clustering on MFCC features to create the initial set of discrete targets, or 'hidden units'.

Step 2: Representation Learning via Masked Prediction

Once we have our target labels (the cluster IDs from Step 1), we can train the main HuBERT model. This step is a direct application of BERT's masked language modeling objective.

  1. The raw audio waveform is passed through the CNN encoder to get a sequence of continuous latent features, let's call them .
  2. A certain percentage of the timesteps in this sequence are masked. This follows the strategy from wav2vec 2.0 and SpanBERT, where whole spans of consecutive timesteps are masked.
  3. This masked sequence is fed into the Transformer encoder, which outputs a sequence of contextualized representations, .
  4. For each masked timestep , the model must predict the corresponding hidden unit (cluster ID) that was generated in Step 1.
  5. A simple cross-entropy loss is calculated between the model's prediction and the target cluster ID.

Crucially, as highlighted in the HuBERT paper, the loss is computed only over the masked regions.

Here, is the set of masked indices, is the masked audio feature sequence, and is the target cluster ID for timestep . This design choice forces the model to learn meaningful representations of the unmasked context to be able to infer the content of the masked regions. The ablation studies in the paper show this is more robust to the noisy nature of the initial clustered labels compared to also calculating loss on unmasked tokens.

Let's watch the final part of the explanation for this step.

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units #nlp

This last segment from the Social Robotics Talk video covers the prediction step and how the cross-entropy loss is applied.

Watch from 12:59 to 17:14. This explains how the raw waveform is processed by the CNN and BERT encoder, and how the model is trained to predict the previously generated cluster targets at the masked positions.


3. Iterative Refinement: Bootstrapping Better Targets

You might be thinking that using MFCCs is a rather basic way to generate targets. While they work for initialization, the representations learned by the Transformer should be far richer. This is where HuBERT's most elegant idea comes in: iterative refinement.

After the first round of training is complete, the entire process is repeated for a second iteration, with one key change:

  • New Target Generation: Instead of extracting MFCCs, we pass the audio data through the HuBERT model trained in the first iteration. We then extract the hidden states from an intermediate Transformer layer (e.g., the 6th layer for the BASE model).
  • Re-clustering: We run K-means clustering on these new, richer features to generate a second, improved set of hidden unit labels.
  • Re-training: A new HuBERT model is then trained from scratch, using this higher-quality set of target labels.

This creates a powerful self-improving loop. The model generates better targets for itself, which in turn allows it to learn even better representations. This bootstrapping process is a core reason for HuBERT's strong performance.

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units #nlp

The concept of iterative refinement is what truly distinguishes HuBERT. This short video segment explains this 'bootstrapping' process.

Watch from 17:14 to 19:22. This crucial part explains how, for the second iteration, the targets are no longer based on MFCCs but on the intermediate representations from the model itself, leading to a cycle of improvement.

For a deeper dive into the methodology and the mathematical formalization, the original paper is invaluable.

[PDF] HuBERT: Self-Supervised Speech Representation Learning by ...

Let's consult the original HuBERT paper to solidify our understanding of the complete methodology.

Read Section II, 'Method' (pages 2-3). This section formally describes the entire process we've just discussed: (A) Learning Hidden Units, (B) Masked Prediction, (D) Iterative Refinement, and (E) Implementation details. This will connect the visual explanations to the underlying research.


4. HuBERT vs. wav2vec 2.0: A Summary

Having explored both models, we can now draw a clear comparison. While their goals and architectures are similar, their pre-training philosophies are distinct.

Feature wav2vec 2.0 HuBERT
Target Generation Online, simultaneous with training. Offline, in a separate, preceding step.
Discretization Method Product Quantization with Gumbel-Softmax. K-means Clustering.
Loss Function Contrastive Loss + Diversity Loss. Cross-Entropy Loss (on masked tokens only).
Target Refinement Learned jointly; no explicit iterative loop. Iterative, using the model's own intermediate features to re-cluster.
Training Stability More complex; requires tuning Gumbel temperature and diversity weight. Simpler and generally more stable.

The HuggingFace team hosted a paper discussion that offers excellent insights into these differences and the advantages of HuBERT's approach.

ML4Audio - HuBERT paper discussion

This video from HuggingFace provides a research-level discussion on HuBERT, comparing its simplified loss function and iterative clustering directly against wav2vec 2.0's approach.

Watch from 39:52 to 43:22. This segment provides a concise recap of HuBERT's objective, highlighting its simplified cross-entropy loss and the two-iteration clustering process, contrasting it with wav2vec 2.0's more complex loss system.


Conclusion

In this lesson, we have thoroughly examined the HuBERT model, a cornerstone of modern self-supervised speech representation learning. You've seen how it cleverly simplifies the pre-training task by separating target discovery from masked prediction, enabling a more stable and familiar BERT-like training objective.

Key Takeaways:

  • Architecture: HuBERT uses the same CNN + Transformer architecture as wav2vec 2.0.
  • Pre-training Objective: It employs a two-step process:
    1. Offline Clustering: Use K-means to discover discrete "hidden units" from audio features.
    2. Masked Prediction: Train a BERT-like model to predict these hidden units at masked positions using a simple cross-entropy loss.
  • Iterative Refinement: The model's key strength lies in its ability to bootstrap. After an initial training run using MFCC-based targets, it uses its own superior intermediate representations to generate better targets for a second training run.
  • Simplicity and Power: By decoupling target generation, HuBERT achieves state-of-the-art results with a simpler, more stable training procedure than its predecessor, wav2vec 2.0.

Preview of the Next Lesson:

We have now covered the pre-training methodologies for both wav2vec 2.0 and HuBERT. The entire purpose of this extensive pre-training on unlabeled data is to create a powerful model that can then be adapted to specific downstream tasks with very little labeled data. In our next lesson, we will put this into practice and learn how to fine-tune a pretrained wav2vec 2.0 model for Automatic Speech Recognition (ASR), and see just how data-efficient this approach can be.

Can't find a good explanation? Sign up and we'll make it for you

Sign up