Hello! Let's dive into the next lesson.
In our previous lesson, we explored autoregressive (AR) vocoders like WaveNet and WaveRNN. We saw how their use of dilated causal convolutions and sample-by-sample generation produced state-of-the-art audio quality. However, we also identified their critical weakness: extremely slow inference due to their sequential nature. This made them impractical for many real-time applications.
Today, we'll examine the powerful solution to this problem: non-autoregressive, GAN-based vocoders. Our focus will be on a landmark model that redefined the trade-off between speed and quality.
The learning outcome for this lesson is to describe the HiFi-GAN vocoder, including its generator and multi-scale/multi-period discriminators. You'll see how a clever adversarial setup can synthesize audio in parallel, achieving both remarkable speed and fidelity that rivals or even surpasses AR models.
1. The Shift to Non-Autoregressive Generation with GANs
The core idea behind GAN-based vocoders is to abandon the sample-by-sample generation process. Instead of modeling the conditional probability , we train a generator network to produce the entire waveform from a mel-spectrogram in a single, parallel forward pass: .
To ensure the generated waveform is indistinguishable from real audio, we introduce a discriminator network . The two networks are trained in opposition:
- The Generator () tries to create waveforms from mel-spectrograms that are so realistic they can fool the discriminator.
- The Discriminator () tries to become an expert at distinguishing real, ground-truth audio from the fake audio produced by the generator.
This adversarial process pushes the generator to learn the intricate details of real audio waveforms. Early models like MelGAN proved this was a viable approach but still left a quality gap compared to WaveNet. HiFi-GAN was the breakthrough that closed this gap.
To get a high-level overview of the problem space and HiFi-GAN's contributions, let's watch a segment from a review video.
Review of HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
This video by Olewave provides excellent context, explaining the role of a vocoder, the limitations of previous models like WaveNet and MelGAN, and introducing the core ideas behind HiFi-GAN.
Please watch from 04:44 to 18:10. Focus on: The explanation of a vocoder's input (mel-spectrogram) and output (waveform). The discussion of why autoregressive models like WaveNet are slow. The introduction of GANs as a faster alternative and the quality limitations of its predecessor, MelGAN.
2. The HiFi-GAN Generator: Multi-Receptive Field Fusion (MRF)
The generator's task is to upsample the low-temporal-resolution mel-spectrogram into a high-resolution waveform. This is primarily done using a stack of transposed convolutions. The key innovation in HiFi-GAN's generator is the Multi-Receptive Field Fusion (MRF) module.
The MRF module is designed to help the generator see patterns of various lengths simultaneously. It consists of multiple parallel residual blocks. Each block contains its own stack of dilated convolutions (a concept you'll remember from WaveNet), but with different kernel sizes and dilation rates. The outputs of these parallel blocks are then summed.
This architecture allows one residual block to focus on short-term patterns (e.g., with small dilations) while another simultaneously focuses on longer-term dependencies (with large dilations).

Let's read the official description from the HiFi-GAN paper to solidify our understanding of the generator and the MRF module.
[PDF] HiFi-GAN: Generative Adversarial Networks for Efficient ...
We'll now turn to the original HiFi-GAN paper by Kong et al. This section formally describes the generator's architecture.
Please read Section 2.2, 'Generator,' and the paragraph 'Multi-Receptive Field Fusion' on page 2. Pay attention to how it describes the role of transposed convolutions and how the MRF module uses multiple residual blocks with different kernel sizes and dilation rates.
3. The Discriminators: Capturing Periodic and Sequential Patterns
The true genius of HiFi-GAN lies in its discriminator design. The authors recognized that a major failure of previous GAN vocoders was their inability to model the subtle periodic patterns present in speech. Human speech is quasi-periodic, composed of fundamental frequencies and harmonics, and our ears are highly sensitive to errors in these patterns, which we perceive as metallic or robotic artifacts.
To address this, HiFi-GAN uses two complementary discriminators: the Multi-Period Discriminator (MPD) and the Multi-Scale Discriminator (MSD).
3.1. Multi-Period Discriminator (MPD)
The MPD is designed specifically to find and critique periodic artifacts. It is a collection of sub-discriminators, each focused on a different periodic pattern.
- Core Idea: Each sub-discriminator is assigned a specific period, (the authors use prime numbers like 2, 3, 5, 7, 11 to minimize redundancy).
- Mechanism: Instead of looking at the 1D waveform directly, the sub-discriminator for period reshapes the input audio. If the audio has length , it's rearranged into a 2D representation of shape .
- Effect: A 2D convolution is then applied to this reshaped data. This structure inherently forces the discriminator to evaluate patterns that repeat every samples. For example, the sub-discriminator with will be sensitive to artifacts that have a 5-sample periodicity.

3.2. Multi-Scale Discriminator (MSD)
While the MPD looks for disjoint periodic patterns, the MSD (an idea borrowed from MelGAN) is used to evaluate the audio structure on a more continuous, sequential basis at different scales.
- Core Idea: The MSD is also a collection of sub-discriminators.
- Mechanism: One sub-discriminator operates on the raw audio directly. The other two operate on versions of the audio that have been downsampled (e.g., by average pooling by a factor of 2 and 4).
- Effect: This allows the MSD to assess the quality of both the fine-grained waveform structure (at the original scale) and the broader, long-term structure (at the downsampled scales).
Let's read the paper's description of this novel discriminator architecture.
[PDF] HiFi-GAN: Generative Adversarial Networks for Efficient ...
Let's return to the HiFi-GAN paper to understand the discriminator design.
Please read Section 2.3, 'Discriminator,' on pages 2-3. Focus on the motivation for modeling periodic patterns and the distinct mechanisms of the 'Multi-Period Discriminator' and 'Multi-Scale Discriminator'.
4. The Loss Functions: A Three-Part Objective
Training this complex system requires a carefully designed set of loss functions that guide the generator toward realism. The total loss for the generator is a weighted sum of three components:
-
GAN Loss (): This is the standard adversarial loss that encourages the generator to create outputs that the discriminators classify as real. HiFi-GAN uses a least-squares formulation (LS-GAN), which tends to be more stable than the original binary cross-entropy loss. The loss is computed for each sub-discriminator in both the MPD and MSD.
-
Mel-Spectrogram Loss (): To ensure the generator produces audio that is a faithful reconstruction of the input condition, an auxiliary L1 loss is added. It measures the absolute difference between the mel-spectrogram of the generated waveform and the original, ground-truth mel-spectrogram.
where is the function that converts a waveform to a mel-spectrogram. -
Feature Matching Loss (): This loss further stabilizes training by comparing the intermediate feature maps of the discriminators. Instead of just fooling the discriminator's final output, the generator is also trained to produce internal representations that are similar to those of real audio. The loss is the L1 distance between the discriminator feature maps for real () and generated () samples, summed across all discriminator layers and all sub-discriminators.
The final generator objective is a weighted sum: .
5. From Theory to Code: HiFi-GAN in SpeechBrain
Your background in CSE and experience with PyTorch frameworks make it valuable to see how these architectural concepts translate to code. The SpeechBrain toolkit provides a clear implementation of HiFi-GAN. You don't need to read this in detail now, but notice the class names and their parameters, which directly map to the concepts we've discussed.
HifiganGenerator: Implements the generator, taking parameters likeupsample_factors,resblock_kernel_sizes, andresblock_dilation_sizesthat define the MRF modules.MultiPeriodDiscriminator: A wrapper class that instantiates multipleDiscriminatorPmodules, each with a different primeperiod.MultiScaleDiscriminator: A wrapper that instantiatesDiscriminatorSmodules, which operate on the original and downsampled audio.HifiganDiscriminator: The top-level class that combines the MPD and MSD into a single unit for training.
This mapping from paper to code is a crucial skill in AI research and development, allowing you to quickly understand and adapt existing models. You can explore the implementation further in the SpeechBrain HifiGAN documentation.
Conclusion
In this lesson, we've broken down the architecture and training strategy of HiFi-GAN, a pivotal model in modern speech synthesis. It solved the speed-quality dilemma that plagued earlier vocoders by introducing a sophisticated GAN architecture.
Key Takeaways:
- HiFi-GAN is a non-autoregressive vocoder that uses a Generative Adversarial Network to synthesize high-fidelity audio in a fast, parallel manner.
- The generator employs a Multi-Receptive Field Fusion (MRF) module, which uses parallel stacks of dilated convolutions to capture audio patterns at multiple scales simultaneously.
- The key innovation is the discriminator, which consists of two distinct components:
- The Multi-Period Discriminator (MPD) specifically targets periodic artifacts by reshaping the audio and using dedicated sub-discriminators for different periods.
- The Multi-Scale Discriminator (MSD) assesses the overall waveform structure at different resolutions (scales).
- Training is stabilized by a combination of an adversarial loss, a mel-spectrogram reconstruction loss, and a feature matching loss.
Preview of the Next Lesson:
Now that you have a deep theoretical understanding of how HiFi-GAN works, it's time to put it into practice. In the next lesson, we will generate an audio waveform from a mel-spectrogram using a pretrained HiFi-GAN model. We will load a model and its corresponding acoustic model (like FastSpeech 2) and run the full inference pipeline to synthesize speech.