Skip to main content
Create your own
Lesson illustration

Autoregressive Vocoders: WaveNet and WaveRNN

Hello! Welcome to the next lesson in our journey through audio AI.

In our previous lesson, we completed our look at the first half of the TTS pipeline—the acoustic model—by comparing the trade-offs between autoregressive (e.g., Tacotron 2) and non-autoregressive (e.g., FastSpeech 2) paradigms. We saw that this was a battle between quality/flexibility and speed/robustness.

Today, we move to the second critical component of a TTS system: the vocoder. This is the neural network responsible for transforming a mel-spectrogram into a high-fidelity, audible waveform. You will see that the same AR vs. NAR battleground exists here, with even more dramatic consequences.

Our goal is to describe autoregressive neural vocoders like WaveNet and WaveRNN, focusing on their groundbreaking use of dilated causal convolutions. These models represented a monumental leap in synthesis quality and laid the groundwork for all modern neural vocoders.


1. The Vocoder's Challenge: From Spectrogram to Waveform

First, let's clarify the vocoder's role. An acoustic model like FastSpeech 2 outputs a mel-spectrogram, a compressed representation of audio. For a 1-second clip sampled at 22,050 Hz, the mel-spectrogram might have around 86 time-steps. The raw audio waveform, however, has 22,050 samples.

The vocoder's task is to perform this massive upsampling, "in-painting" the phase information and high-frequency details lost in the spectrogram to generate a realistic waveform. For decades, this was done with signal processing algorithms like Griffin-Lim, which often produced robotic or artifact-laden audio. Neural vocoders changed the game by learning to synthesize audio directly.

2. WaveNet: A Generative Model for Raw Audio

The major breakthrough came in 2016 from DeepMind with WaveNet. Its approach was radical at the time: model the probability distribution of the raw audio waveform directly.

WaveNet is a fully autoregressive model. It formulates the joint probability of a waveform as a product of conditional probabilities:

This means that to generate the audio sample at timestep , the model is conditioned on all the samples that came before it. It predicts the audio one sample at a time.

To understand the core ideas directly from the source, let's read the introduction and problem formulation from the original paper.

wavenet:agenerative model for raw audio

We'll start with the original WaveNet paper from DeepMind. This will introduce the model's core autoregressive concept.

Please read the 'ABSTRACT' and the first part of Section 2, 'WAVENET,' up to the formula. Focus on how it defines the task as generating raw audio by modeling the conditional probability of each sample.

A crucial detail mentioned in the paper is that modeling 16-bit audio directly would require a softmax output of 65,536 probabilities per sample. To make this tractable, WaveNet first applies µ-law companding, a non-linear quantization technique that transforms the audio into 256 discrete values. This turns the regression problem into a much simpler 256-class classification problem.

3. The Core Innovation: Dilated Causal Convolutions

The autoregressive formulation presents a huge challenge: for audio at 24 kHz, a single second of sound contains 24,000 samples. To generate a coherent sound, the model's prediction at sample 24,000 might need to depend on information from sample 1. How can a model have such a massive receptive field?

Recurrent Neural Networks (RNNs) were the standard for sequences, but they are notoriously difficult to train on very long sequences due to vanishing gradients and are slow due to their sequential nature. WaveNet's solution was to use convolutions, but with two special properties.

Let's watch a video that provides an excellent conceptual breakdown of this idea, which is often called a Temporal Convolutional Network (TCN).

Lecture 5.4 - CNNs for Sequential Data

This video from the 'DLVU' channel clearly explains the concepts of causal and dilated convolutions, which are the building blocks of WaveNet.

Please watch the segment from 01:42 to 12:30. Pay close attention to: Causal Convolutions (01:42 - 04:44): How asymmetric padding ensures the model doesn't 'see' the future. Dilated Convolutions (04:44 - 09:05): How 'stretching' the convolution kernel allows the receptive field to grow without increasing parameters. Stacking Layers (09:05 - 12:30): How stacking layers with exponentially increasing dilation factors creates a very large receptive field efficiently.

Let's formalize these two concepts.

3.1. Causal Convolutions

To maintain the autoregressive property—that a prediction cannot depend on future samples —the model must be causal. For a 1D convolution, this is achieved by padding the input sequence only at the beginning (the "left" side). This ensures that the filter at any given timestep can only see current and past inputs.

Dilated Causal Convolution Architecture
A visualization of stacked dilated causal convolutions. Notice how the output at any position only depends on inputs to its left (from the past). The 'dilation' causes the receptive field to expand rapidly, covering all 16 inputs with just 4 layers.

3.2. Dilated Convolutions

To efficiently increase the receptive field, WaveNet uses dilated convolutions. Instead of the filter being applied to adjacent samples, it's applied to samples with a certain step size, or "dilation rate."

By stacking these layers and doubling the dilation rate at each layer (e.g., 1, 2, 4, 8, 16, ...), the receptive field grows exponentially with depth. This allows the model to capture dependencies across thousands of timesteps with a relatively small number of layers, solving the long-range dependency problem far more efficiently than standard convolutions or RNNs.

Now, let's quickly review these concepts in the original paper for their formal definition.

wavenet:agenerative model for raw audio

Let's return to the WaveNet paper to see how the authors formally describe causal and dilated convolutions.

Read Section 2.1, 'DILATED CAUSAL CONVOLUTIONS.' This section explains both concepts and how they are combined. The figures are particularly helpful for visualizing the structure.

4. The Full WaveNet Architecture

The dilated causal convolutions are the core engine, but they are assembled into a sophisticated architecture using a few more key components.

WaveNet Dilated Causal Convolution Block Architecture
A diagram of a single WaveNet residual block. It shows the flow of information through the dilated convolution, the gated activation (tanh and sigmoid), and the creation of the residual and skip connections.

Let's watch a short video that walks through this architectural block.

WaveNet (Theory and Implementation)

This video from CanConTech provides a great walkthrough of the WaveNet architectural block and the role of each component.

Watch the segment from 04:54 to 06:55. Focus on how the input is processed by the gated activation and how the residual and skip connections are produced.

As shown in the diagram and video, each block in the network consists of:

  1. Gated Activation Unit: Instead of a standard ReLU, WaveNet uses a gated activation:

    Here, the input is passed through two separate convolutional layers. One output is passed through a function (the "filter") and the other through a sigmoid function (the "gate"). The gate multiplicatively controls which information from the filter is passed forward. This structure has been found to be more effective for modeling complex data like audio.

  2. Residual and Skip Connections: The output of the gated activation is processed by a 1x1 convolution to produce two paths:

    • A residual connection, which is added back to the block's input and passed to the next block. This helps in training very deep networks by preventing gradient vanishing.
    • A skip connection, which is routed to a final summation block. The outputs from all blocks are summed, passed through some final processing layers (ReLUs and 1x1 convolutions), and finally to a softmax layer to predict the next audio sample. This allows the final prediction to benefit from features at all levels of abstraction in the network.
  3. Local Conditioning: To be used as a vocoder, the network needs to be conditioned on the mel-spectrogram. This is done by first upsampling the mel-spectrogram to match the high temporal resolution of the audio. This upsampled representation is then added to the input of the gated activation unit inside each block, guiding the audio generation process.

5. WaveRNN: A More Compact Alternative

WaveNet's convolutional structure, while powerful, was computationally intensive during inference. Because generation is sample-by-sample, the entire large network must be evaluated for every single sample.

WaveRNN was proposed as a more compact and efficient autoregressive alternative.

  • Core Idea: Replace the stack of large convolutions with a much smaller Recurrent Neural Network (typically a single-layer GRU).
  • Challenge: The matrix multiplications in a standard RNN would still be too slow for sample-level generation.
  • Solution: WaveRNN uses several optimization tricks. The most important is decomposing the prediction: instead of one 16-bit prediction, it predicts the 8 most significant bits (the "coarse" part) first, and then predicts the 8 least significant bits (the "fine" part) conditioned on the coarse part. It also leverages weight pruning to create extremely sparse matrices, which can be computed very efficiently, especially on CPUs.

While still autoregressive and thus slower than NAR models, WaveRNN provided a significant speed-up over the original WaveNet, making real-time AR vocoding on a CPU feasible for the first time.

Conclusion

In this lesson, we dissected the architecture of pioneering autoregressive vocoders. These models set a new standard for audio quality by directly modeling the raw waveform.

Key Takeaways:

  • Autoregressive vocoders generate audio one sample at a time, conditioning each new sample on all previous ones.
  • WaveNet's key innovation was the use of dilated causal convolutions, which allow the network's receptive field to grow exponentially with depth. This efficiently captures the long-range temporal dependencies crucial for realistic audio.
  • The full WaveNet architecture combines these convolutions with gated activation units and residual/skip connections to enable stable training of deep, powerful models.
  • WaveRNN offered a more computationally efficient alternative by using a compact RNN with clever optimization tricks.
  • The primary drawback of all AR vocoders is their slow, sequential inference, which motivated the search for faster parallel models.

Preview of the Next Lesson:

The extreme slowness of AR vocoders was a major bottleneck for real-world applications. In our next lesson, we will explore the solution: non-autoregressive vocoders. We will focus on HiFi-GAN, a model that leverages Generative Adversarial Networks to achieve both state-of-the-art audio quality and incredibly fast parallel synthesis, solving the speed problem that plagued WaveNet.

Can't find a good explanation? Sign up and we'll make it for you

Sign up