Hello! Welcome to your next lesson in the Text-to-Speech Architectures module.
In our last session, we outlined the standard neural TTS pipeline, breaking it down into a frontend (text processing) and a backend (audio synthesis). We established that the backend typically consists of two key components: an acoustic model that generates a mel-spectrogram from text, and a vocoder that synthesizes a waveform from that spectrogram.
Today, we will perform a deep dive into one of the most influential acoustic models: Tacotron 2. Our goal is to satisfy the learning outcome: Describe the Tacotron 2 architecture, focusing on its encoder, location-sensitive attention, and autoregressive decoder. We will dissect the model piece by piece, understanding not just what each component is, but why it was designed that way.
By the end of this lesson, you will have a detailed architectural understanding of how Tacotron 2 transforms a sequence of characters into a detailed mel-spectrogram, ready for synthesis.
1. Tacotron 2: A High-Level Overview
Tacotron 2, introduced by Google in 2017, represented a major leap forward in speech synthesis quality. It is a sequence-to-sequence model that forms the "acoustic model" part of the two-stage TTS pipeline.
- Input: A sequence of characters. (One of its key innovations was working directly on characters, often bypassing the need for an explicit grapheme-to-phoneme conversion step).
- Output: A mel-spectrogram.
This generated mel-spectrogram is then passed to a neural vocoder (like WaveNet, which was used in the original paper, or a more modern vocoder like HiFi-GAN) to produce the final audio waveform.
Let's start with a brief video that positions Tacotron 2 in the context of the neural TTS revolution.
Text-to-Speech & Voice Cloning Course: Neural TTS Revolution
This clip from Valerio Velardo provides a concise introduction to Tacotron 2, highlighting its role as a breakthrough acoustic model and its key innovations.
Watch the segment from 00:19:38 to 00:26:50. Pay attention to how Tacotron 2 fits into the two-stage pipeline, the mention of its sequence-to-sequence nature, and the critical role of the attention mechanism.
As the video explains, Tacotron 2 combines several powerful deep learning concepts into a unified, end-to-end trainable system. The architecture can be broken down into three main parts, which we will explore in detail:
- An Encoder that creates a representation of the input text.
- An Attention Mechanism that aligns the text with the audio.
- A Decoder that generates the mel-spectrogram frame by frame.
Let's look at the complete architecture diagram to get a bird's-eye view before we dive into the components.

Now, let's unpack each of these modules.
2. The Encoder: Capturing Textual Context
The encoder's job is to take the input character sequence and convert it into a sequence of hidden states, or "annotations," that are rich with contextual information.
Spectrogram Feature prediction network
This GitHub wiki page for a Tacotron 2 implementation provides an excellent, technically detailed breakdown of the model. We'll start with the section on the encoder.
Read the section titled 'Encoder'. Focus on understanding the two main parts of the encoder: the stack of convolutional layers and the bidirectional LSTM. Note the rationale for using this combination.
Let's summarize the key aspects of the encoder design:
-
Input: The process starts with a sequence of input characters, which are first passed through a character embedding layer to get a continuous vector representation for each character.
-
Convolutional Stack: The character embeddings are then processed by a stack of three 1D convolutional layers. Given your background in deep learning, you'll recognize this pattern. Here, the convolutional filters act as n-gram detectors, learning to extract local, contextual features from the character sequence. For example, they can learn that the character sequence "t-i-o-n" often corresponds to a "shun" sound. This helps the model capture pronunciation nuances that depend on neighboring characters.
-
Bidirectional LSTM: The feature maps from the convolutional stack are fed into a single bidirectional LSTM layer. Since you're familiar with recurrent architectures, you know that the bi-LSTM processes the sequence both forwards and backwards. This allows each encoder hidden state to contain information about the entire input sentence—both the characters that came before and those that come after. This global context is crucial for determining correct pronunciation and prosody.
The final output of the encoder is a sequence of annotation vectors , where is the length of the input character sequence. Each vector is a rich representation of the j-th character in the context of the full sentence.
3. Location-Sensitive Attention: Aligning Text and Speech
Now we come to one of the most critical innovations in Tacotron 2. The attention mechanism is the bridge between the encoder and the decoder. At each step of generating the spectrogram, it must decide which part of the encoded text to "focus on" or "attend to."
Speech is monotonic; we say words in order. A standard content-based attention mechanism (like the one used in machine translation) can get confused by repetitive sounds or long silences, causing it to jump around, repeat words, or skip parts of the text. To solve this, Tacotron 2 uses Location-Sensitive Attention.
This mechanism enhances standard additive attention by making it aware of the alignment decisions from previous steps. This encourages the attention to move forward consistently from left to right through the input text.
Spectrogram Feature prediction network
Let's return to the GitHub wiki page to understand how this attention mechanism works. It provides a great step-by-step evolution from content-based to location-sensitive attention.
Read the section 'Attention Mechanism'. Follow the progression from 'Content based attention' to 'Location based attention' and finally to 'Tacotron-2 Attention'. Focus on how cumulative attention weights are used to create 'location features'.
Let's break down the core idea:
- Standard Additive Attention: At each decoder step
i, an alignment score is calculated between the decoder state and each encoder state . - Location-Sensitive Addition: Tacotron 2 adds a new term to this equation. It first calculates location features by passing the cumulative attention weights from all previous steps, , through a 1D convolutional layer. This essentially summarizes "where the attention has been so far."
- The New Score: This location feature is then incorporated into the energy calculation. The new score now depends on the decoder state, the encoder state, and the location features. (Note: The indices and inputs can vary slightly between implementations, but the core idea of adding a location-dependent term remains).
By adding this location-based term, the model is explicitly penalized for attending to the same places repeatedly. It learns to shift its attention smoothly and monotonically forward through the input sequence, which is exactly how human speech is produced.
For more intuition, you can optionally skim this blog post which provides a nice summary and links the theory to PyTorch code.
Location based attention – from “Attention based Models for Speech ...
This blog post offers another perspective on location-based attention, referencing the original research paper by Chorowski et al. and showing snippets of a PyTorch implementation.
(Optional) Skim this post to see how the mathematical concept of convolving previous alignments is translated into code. The code snippets under the 'Tacotron2' heading are particularly relevant.
4. The Autoregressive Decoder and Post-Net
The decoder's task is to generate the mel-spectrogram. It's an autoregressive model, which means that to generate the output for the current time step, it uses the output it generated in the previous time step as input.
Let's walk through one decoding step, referring to the diagram above.
Spectrogram Feature prediction network
The final component is the decoder. The same wiki page provides a clear, step-by-step description of the decoding loop.
Read the section 'Decoder'. Pay close attention to the flow of information: the role of the Pre-Net, the 'input feeding' approach, how the attention context vector is used, and the two final outputs (spectrogram frame and stop token).
Here is a summary of the process at each decoder time step i:
-
Pre-Net: The mel-spectrogram frame from the previous step, , is passed through a "Pre-Net"—a small stack of fully-connected layers with dropout. The paper's authors describe this as an information bottleneck that helps stabilize training and improve generalization.
-
Decoder RNN: The output from the Pre-Net is concatenated with the attention context vector from the previous step, . This is called input-feeding. It informs the decoder about the alignment from the previous step before it makes its next prediction. This combined vector is fed into a stack of two unidirectional LSTMs.
-
Attention Query: The output of the decoder LSTMs, , is used as the "query" for the location-sensitive attention mechanism. The attention mechanism uses this query and the encoder outputs to compute the new context vector, .
-
Prediction: The LSTM output is concatenated with the new context vector . This final vector is passed through two separate linear projection layers to predict:
- The next mel-spectrogram frame, . The model can be configured to predict
rframes at a time (whereris the reduction factor) to speed up generation. - A stop token probability. This is a single scalar value passed through a sigmoid function. When this probability crosses a certain threshold, the decoding process stops. This allows the model to dynamically decide when the utterance is finished.
- The next mel-spectrogram frame, . The model can be configured to predict
The Post-Net
The spectrogram generated directly by the decoder can sometimes be blurry or lack fine detail. To fix this, Tacotron 2 adds a Post-Net.
- Architecture: A stack of five 1D convolutional layers.
- Function: The Post-Net does not predict the spectrogram directly. Instead, it takes the "coarse" spectrogram from the decoder as input and predicts a residual. This residual is then added back to the coarse spectrogram to create a final, refined version.
This two-stage prediction (coarse + residual) is a common technique in signal processing models, as it's often easier for a network to learn to correct small errors than to predict a perfect output from scratch. The entire model is trained with an L2 loss on both the coarse and the refined spectrograms, plus a binary cross-entropy loss for the stop token.
Conclusion
In this lesson, we have conducted a thorough architectural review of Tacotron 2, a landmark model in neural speech synthesis. We've seen how it elegantly combines several key deep learning components to map text to a mel-spectrogram.
Key Takeaways:
- Overall Structure: Tacotron 2 is a sequence-to-sequence acoustic model with an encoder-decoder architecture, bridged by an attention mechanism.
- Encoder: Uses a CNN + bi-LSTM stack to produce a sequence of contextual character representations. The CNNs capture local features (like n-grams), while the bi-LSTM captures global sentence-level context.
- Location-Sensitive Attention: This is the key to monotonic alignment in speech. By incorporating cumulative attention weights from previous steps as a feature, it learns to move forward through the text, avoiding repetitions and skips.
- Autoregressive Decoder: Generates the mel-spectrogram one frame at a time. It uses a Pre-Net as an information bottleneck and input-feeding to provide the decoder with its previous alignment decision. It simultaneously predicts a stop token to end generation dynamically.
- Post-Net: A final convolutional stack that predicts a residual to refine the coarse spectrogram from the decoder, improving overall audio quality.
Preview of the Next Lesson:
A defining characteristic of Tacotron 2 is its autoregressive nature. While this allows it to generate very high-quality output, it also makes inference slow, as frames must be generated sequentially. In our next lesson, we will explore FastSpeech 2, a non-autoregressive TTS model that addresses this limitation by generating all spectrogram frames in parallel. We will compare its architecture to Tacotron 2 and analyze the trade-offs between speed, quality, and controllability.