Skip to main content
Create your own
Lesson illustration

Autoregressive vs. Non-Autoregressive TTS: A Comparative Analysis

Hello! Welcome back to our course on audio AI.

In the last two lessons, we explored two landmark Text-to-Speech architectures. First, we dissected Tacotron 2, a powerful autoregressive (AR) model known for its high-quality output generated sequentially. Then, we investigated FastSpeech 2, a pioneering non-autoregressive (NAR) model that achieves massive speed gains by generating speech in parallel.

Today, we'll place these two paradigms side-by-side. Our goal is to compare autoregressive vs. non-autoregressive TTS models in terms of synthesis quality, speed, and controllability. This comparison isn't just about two models; it's about understanding a fundamental trade-off that has shaped the entire field of generative modeling for speech. By the end, you'll have a clear framework for evaluating TTS systems and understanding their respective strengths and weaknesses.


1. The Core Architectural Divide: Sequential vs. Parallel

Let's start by reinforcing the fundamental difference in how these models generate output.

Autoregressive vs. Non-Autoregressive Models
A visual comparison between (a) an autoregressive model, where the prediction at each step depends on the previous output, and (b) a non-autoregressive model, where all outputs are predicted independently and in parallel.

Autoregressive (AR) Models

An AR model factorizes the probability of an output sequence given an input as a chain of conditional probabilities:

In TTS, this means each mel-spectrogram frame is generated conditioned on all previously generated frames .

  • Example: Tacotron 2. Its decoder uses an attention mechanism to read the input text and combines that with the previously generated frame to produce the next frame. This process repeats until a "stop token" is predicted. The sequential nature is the source of its quality but also its primary bottleneck.

The following short video, though focused on machine translation, perfectly explains why this autoregressive decoding process is inherently slow.

Non-Autoregressive and Shallow Decoding: Speeding up Translation

This clip from the Efficient NLP channel clearly explains the computational bottleneck of autoregressive decoding.

Watch the segment from 00:39 to 01:20. Notice how the number of sequential steps is tied to the length of the output sequence, making it a slow process.

Non-Autoregressive (NAR) Models

An NAR model simplifies the problem by assuming conditional independence among the output elements given the input:

This allows the model to generate all frames of the mel-spectrogram simultaneously in a single forward pass.

  • Example: FastSpeech 2. It uses a duration predictor to determine the length of the output sequence upfront and then generates the entire mel-spectrogram in parallel.

This parallelism promises huge speed-ups, but it introduces a major challenge: the multimodality problem. A single sentence can be spoken in many ways (different pitches, rhythms, etc.). AR models navigate this by continuing a specific prosodic path they've started. NAR models, generating everything at once, risk averaging all possibilities, leading to bland speech. They require explicit mechanisms, like FastSpeech 2's Variance Adaptor, to handle this.


2. A Head-to-Head Comparison

Let's systematically compare AR and NAR models across four critical dimensions: speed, quality, robustness, and controllability.

A. Synthesis Speed

  • Autoregressive (AR): Slow. The number of sequential operations is proportional to the length of the output audio. To generate one second of audio (e.g., ~86 mel-spectrogram frames at a hop length of 256 and a sample rate of 22050 Hz), the model must perform ~86 sequential decoding steps. This makes them poorly suited for real-time, low-latency applications.
  • Non-Autoregressive (NAR): Extremely Fast. The entire output is generated in a single feed-forward pass, making the inference time independent of the audio length. The speedup can be one or two orders of magnitude.

To see some concrete numbers, let's look at the results from the FastSpeech 2 paper.

FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

We'll consult the FastSpeech 2 article from Microsoft Research, which provides a direct comparison of performance metrics.

Focus on 'Table 2: The comparison of training time and inference latency in waveform synthesis.' Compare the 'Inference Latency' of Transformer TTS (an AR model) with FastSpeech 2 (an NAR model). The 'RTF' (Real-Time Factor) is particularly illustrative.

As you saw in Table 2, FastSpeech 2 achieves a waveform synthesis speedup of over 47x compared to an AR Transformer TTS model. This is the primary motivation for developing NAR models.

B. Synthesis Quality

  • Autoregressive (AR): Historically Higher Quality. The step-by-step generation process is excellent at capturing fine-grained details and local dependencies in the audio, often leading to more natural and smooth-sounding speech. The attention mechanism provides a powerful, flexible way to learn text-to-speech alignment.
  • Non-Autoregressive (NAR): Historically Lower Quality. Early models struggled with the multimodality problem, often producing "over-smoothed" speech that lacked prosodic variation. Errors in the explicit duration alignment could also degrade quality.

However, this gap has narrowed dramatically. With better alignment methods (like using pre-trained forced aligners) and sophisticated conditioning (like FastSpeech 2's Variance Adaptor), modern NAR models now achieve quality on par with their AR counterparts.

Let's check the evidence from the FastSpeech 2 article again.

FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Let's look back at the FastSpeech 2 article, this time focusing on audio quality.

This time, examine 'Table 1: The MOS evaluation.' Compare the Mean Opinion Score (MOS) of FastSpeech 2 with Tacotron 2 and Transformer TTS. A higher MOS indicates better perceived quality.

The MOS scores show that FastSpeech 2 is competitive with, and in their experiment slightly better than, the AR models, demonstrating that the quality gap is no longer a major concern for well-designed NAR systems.

C. Robustness and Stability

  • Autoregressive (AR): Less Robust. The sequential feedback loop means that errors can accumulate. A small mistake early in generation can throw off the rest of the sequence. This leads to common failure modes like repeating words, skipping words, or babbling nonsensically, especially with out-of-distribution text. These attention-related failures make AR models less reliable for production use.
  • Non-Autoregressive (NAR): Highly Robust. The feed-forward structure and explicit duration modeling make these models very stable. They don't suffer from the same repetition or skipping issues because the output structure is fixed by the duration predictor. This stability is a significant advantage for real-world deployment.

D. Controllability

  • Autoregressive (AR): Poor Controllability. Prosody (pitch, duration, energy) is learned implicitly as part of the sequence generation process. It is difficult to directly manipulate these attributes at inference time without complex architectural changes.
  • Non-Autoregressive (NAR): Excellent Controllability. Models like FastSpeech 2 that use a Variance Adaptor explicitly model these prosodic features. As a result, they can be easily manipulated. You can adjust the output of the duration predictor to change the speech rate, or modify the pitch predictor's output to alter the intonation of the synthesized voice.

The FastSpeech 2 article provides a clear demonstration of this.

FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Finally, let's see the evidence for controllability in the FastSpeech 2 article.

Read the subsection 'Variance Control' and look at 'Figure 2'. This figure shows how the pitch contour (F0) of the synthesized speech can be directly manipulated by changing the pitch predictor's input.


3. The Big Picture: A "Battle" of Paradigms

The trade-offs between AR and NAR models have led to a dynamic evolution in the field of speech synthesis. To get a comprehensive overview of this "battle," let's study a few slides from a presentation by Xu Tan, one of the authors of the FastSpeech papers.

[PDF] autoregressive Battle in Speech Synthesis - Xu Tan

These slides from Xu Tan's presentation, 'The Battle Between AR and NAR in Neural TTS', provide an expert summary of the state of the field, the pros and cons of each approach, and the trends over time.

Please study the following slides from the PDF: Slide 21 ('Lesson 6: The AR/NAR Battle Is Not A Zero-Sum Game'): Focus on the 'Pros' and 'Cons' table. This is a perfect summary of our comparison. Slide 6-7 ('The Battle Between AR and NAR'): Look at the timeline. Notice how early models were AR (Tacotron), followed by a wave of NAR models (FastSpeech), and how both paradigms continue to evolve. Slide 22 ('Lesson 6: The AR/NAR Battle Is Not A Zero-Sum Game'): Read the bullet points. They emphasize that each model type has different ideal application scenarios.

Let's consolidate the key points from that reading into a summary table.

Feature Autoregressive (AR) Models Non-Autoregressive (NAR) Models
Speed Slow (sequential generation) Fast (parallel generation)
Quality High, very natural prosody High, but can be "over-smoothed" without proper conditioning
Robustness Prone to errors (repeating, skipping) Highly stable and robust
Controllability Poor (implicit prosody) Excellent (explicit modeling of pitch, duration, energy)
Example Tacotron 2, Transformer TTS FastSpeech 2, Glow-TTS
Best For High-quality offline synthesis, research on expressiveness Real-time applications, controllable synthesis, production systems

As Xu Tan's slides highlight, this isn't a zero-sum game. The development of large language models (LLMs) has revived interest in AR models (like VALL-E, which we'll see in Module 10) for their incredible in-context learning and zero-shot synthesis capabilities, where inference speed is less critical than flexibility. At the same time, NAR models have become the standard for efficient, controllable, and robust TTS systems and have come to dominate the vocoder space.


Conclusion

Today, we systematically compared the two dominant paradigms in acoustic modeling for TTS: autoregressive and non-autoregressive generation.

Key Takeaways:

  • AR models (e.g., Tacotron 2) generate speech sequentially, which historically offered higher quality and naturalness but at the cost of slow inference speeds and lower robustness.
  • NAR models (e.g., FastSpeech 2) generate speech in parallel, providing massive speedups, excellent robustness, and fine-grained controllability, making them ideal for production systems.
  • The choice between AR and NAR is a trade-off between speed/robustness and (historically) quality/flexibility.
  • Modern NAR models have largely closed the quality gap through innovations like the Variance Adaptor.
  • The "best" paradigm depends on the application: NAR for low-latency production, and large-scale AR for cutting-edge zero-shot research.

Preview of the Next Lesson:

So far, we've focused on the first stage of the TTS pipeline: the acoustic model that converts text to mel-spectrograms. In our next lesson, we will shift our focus to the second stage: the vocoder, which converts mel-spectrograms into audible waveforms. We will start by exploring two foundational autoregressive vocoders: WaveNet and WaveRNN. You will see that the AR/NAR distinction is just as critical in this domain, with its own set of fascinating trade-offs.

Can't find a good explanation? Sign up and we'll make it for you

Sign up