Skip to main content
Create your own
Lesson illustration

Speech Recognition: Sequence-to-Sequence and Alignment Challenges

Hello! Welcome to the sixth module of our course. In the previous module, we explored the foundations of deep generative modeling, comparing the strengths and weaknesses of VAEs, GANs, and Flow-based models. We're now going to pivot from unsupervised generation to a core supervised task in audio AI: Automatic Speech Recognition (ASR).

Your experience with models like Whisper has given you a practical sense of what ASR can do. Our goal now is to peel back the layers and understand the fundamental problem these models are designed to solve.

Today's lesson focuses on the learning outcome: Formulate speech recognition as a sequence-to-sequence problem and explain the alignment challenge. We will define the ASR task formally, understand why it's so difficult, and introduce the two main conceptual frameworks—Connectionist Temporal Classification (CTC) and attention mechanisms—that form the basis of modern solutions.

1. From Complex Pipelines to End-to-End Systems

Modern ASR systems that map raw audio directly to text are often called "end-to-end" models. To appreciate why this is a significant development, it's helpful to first understand the classical approach.

Lecture 12: End-to-End Models for Speech Processing

Let's begin with a video from Stanford's "End-to-End Models for Speech Processing" lecture series. This first clip will provide a basic definition of ASR and then describe the traditional, multi-component pipeline.

Please watch from 01:54 to 09:30. As you watch, focus on identifying the three main components of the classical ASR pipeline and the motivation for replacing this complex system with a single, end-to-end model.

As the video explained, traditional ASR systems are modular pipelines, typically consisting of:

  1. Acoustic Model (AM): Maps short audio frames (features like MFCCs) to probabilities of phonetic units (phonemes). This is where a neural network was often first introduced to replace older methods like Gaussian Mixture Models (GMMs).
  2. Pronunciation Model (or Lexicon): A dictionary that maps sequences of phonemes to words. For example, it knows that the phoneme sequence /k/ /æ/ /t/ corresponds to the word "cat".
  3. Language Model (LM): Assigns probabilities to sequences of words (e.g., "how are you" is more probable than "how are shoe"). This helps resolve ambiguity and correct errors from the acoustic model.
Traditional ASR Pipeline: From Transcript to Acoustic Features
A diagram illustrating the traditional ASR pipeline. It involves separate models for acoustics, pronunciation, and language, each requiring expert knowledge and separate training. Source: Mael Fabien.

The key takeaway is that this traditional approach required hand-engineered components, linguistic expertise (for phonemes), and careful, separate tuning of each part. The "neural network invasion" replaced these parts one by one, which naturally led to the next question: Can we replace the entire pipeline with a single, unified neural network?

This is the core idea of end-to-end ASR: learning a single function that maps a sequence of audio inputs, , directly to a sequence of text outputs, . We want to model the conditional probability .

2. ASR as a Sequence-to-Sequence (Seq2Seq) Problem

This "sequence of inputs to sequence of outputs" formulation is a general machine learning paradigm known as a sequence-to-sequence (Seq2Seq) problem. It's not unique to speech; machine translation ("Hello, how are you?" → "¿Hola, cómo estás?") is the classic example.

Let's watch a brief, clear explanation of Seq2Seq models.

Sequence-to-Sequence (seq2seq) Encoder-Decoder Neural Networks, Clearly Explained!!!

The StatQuest video "Sequence-to-Sequence (seq2seq) Encoder-Decoder Neural Networks" provides an excellent, intuitive overview of the problem and the general architecture used to solve it.

Watch from 01:10 to 03:44. Pay close attention to the core challenge identified: the input and output sequences can have different lengths.

As you just saw, the defining characteristic of a Seq2Seq problem is that the model must handle variable-length inputs and outputs, and the lengths don't have to match. This perfectly describes ASR:

  • Input Sequence : A sequence of feature vectors, , where each is a frame of audio features (e.g., a slice of a mel-spectrogram). The length can be thousands of frames.
  • Output Sequence : A sequence of tokens, , where each is a character or a word. The length is typically much smaller than .

The standard architecture for solving Seq2Seq problems is the Encoder-Decoder model.

Basic RNN Encoder-Decoder Architecture
A conceptual diagram of an Encoder-Decoder model. The encoder processes the entire input sequence and compresses it into a context vector. The decoder then uses this context vector to generate the output sequence one step at a time. Source: Mael Fabien.

The architecture works as follows:

  1. Encoder: A neural network (often a Recurrent Neural Network like an LSTM, or more recently, a Transformer) processes the entire input sequence and compresses its information into a fixed-size representation called a context vector. This vector is expected to be a meaningful summary of the entire audio input.
  2. Decoder: Another neural network takes the context vector from the encoder as its initial state. It then generates the output sequence one token at a time, with each generated token influencing the next.

This architecture elegantly decouples the input and output, allowing their lengths to be different. However, for ASR, there's a deeper, more fundamental challenge that this simple picture doesn't fully capture.

3. The Alignment Challenge

The core difficulty in ASR is that we don't know the alignment between the input audio frames and the output text characters. When you have a 10-second audio clip and a 50-character transcript, which specific milliseconds of audio correspond to the first character? Which correspond to the second? The answer is not obvious and is the central problem that ASR models must solve.

Let's read a concise and clear description of this problem.

Sequence Modeling with CTC

The article "Sequence Modeling with CTC" on Distill.pub provides one of the best explanations of the alignment problem in the context of speech recognition.

Please read the "Introduction" section of the article. Focus on how it frames the problem: we have audio clips and transcripts, but we don't know how the characters align to the audio.

The article you just read, along with this excellent summary from another resource, highlights the key facets of the alignment challenge:

Audio Deep Learning Made Simple - ASR: How it Works

The blog post "Audio Deep Learning Made Simple - ASR: How it Works" also has a fantastic, visually-supported section explaining this issue.

Please read the section titled "Align the sequences". It breaks down the problem into several concrete points.

To summarize, the alignment is difficult because:

  • Length Mismatch: The input sequence (audio frames) is much longer than the output sequence (characters). A single spoken character can span multiple audio frames.
  • No Obvious Boundaries: Speech is a continuous signal. There are no clear markers in the waveform that say "character 'c' ends here, character 'a' begins here."
  • Rate Variation: People speak at different speeds. The duration of a phoneme for a given character can vary significantly.
  • Repeated Characters vs. Elongation: In the word "hello", the audio for the 'l' sound is elongated. How does a model know to output "ll" instead of just a single "l"?
  • Pauses and Non-speech: Audio contains silence, breaths, and filler words ("um", "uh"). The model must learn to ignore or filter these out, producing no corresponding output characters.

Manually aligning every character in a large dataset is prohibitively expensive. Therefore, the most successful ASR models are those that can learn this alignment automatically.

4. Two Primary Approaches to Automatic Alignment

Modern end-to-end ASR models generally use one of two main strategies to solve the alignment problem.

Approach 1: Connectionist Temporal Classification (CTC)

CTC provides an ingenious way to train an ASR model without an explicit alignment.

Lecture 12: End-to-End Models for Speech Processing

Let's return to the Stanford lecture for a high-level overview of Connectionist Temporal Classification (CTC).

Watch the segment from 11:54 to 17:18. Don't worry about the dynamic programming details for now. Focus on understanding the role of the special 'blank' token and the overall idea of collapsing a long sequence of predictions into the final, shorter transcript.

The core ideas behind CTC are:

  1. Frame-wise Prediction: The model (typically an RNN or Transformer) predicts a character probability distribution for every single input frame. This results in a very long sequence of predictions, equal in length to the input audio frames ().
  2. The Blank Token (ϵ): The output vocabulary is augmented with a special blank token. This token represents "not a character" and can be emitted for frames corresponding to pauses or the transitions between characters.
  3. The Collapse Function: A simple set of rules defines how to map the long prediction sequence to the final output:
    • First, merge all consecutive repeated characters (e.g., h, h, e, l, l, l, o → h, e, l, o).
    • Second, remove all blank tokens (e.g., ϵ, c, ϵ, a, a, t, ϵ → c, a, a, t → c, a, t).

The blank token is crucial for handling repeated characters. An alignment for "hello" must have a blank between the two 'l's (e.g., h, e, l, ϵ, l, o) to prevent them from being collapsed into a single 'l'.

CTC works by calculating a loss function that sums the probabilities of all possible valid alignments that could collapse to the correct transcript. This allows the model to learn without ever needing to commit to a single, hard alignment.

Approach 2: Attention-Based Encoder-Decoder

The second approach tackles alignment in a more direct, but still learned, fashion using an attention mechanism.

Lecture 12: End-to-End Models for Speech Processing

This final clip from the Stanford lecture introduces the Listen, Attend, and Spell (LAS) model, a classic example of an attention-based ASR system.

Watch from 26:05 to 30:55, and then from 31:20 to 34:57. The key idea to grasp is how the decoder, at each step of generating an output character, 'attends' to different parts of the encoded audio input.

In an attention-based model:

  1. The Encoder (the "Listen" part) processes the entire audio sequence to produce a set of rich hidden states, one for each input frame.
  2. The Decoder (the "Spell" part) generates the output transcript one character at a time.
  3. At each step, before producing the next character, the decoder uses an Attention Mechanism (the "Attend" part). It compares its current internal state to all of the encoder's hidden states to compute an "attention vector". This vector assigns high weights to the audio frames that are most relevant for predicting the next character.
  4. The decoder then uses a weighted average of the encoder states (weighted by the attention vector) to inform its prediction.

In essence, the model learns to align by itself. As it generates the transcript, the "spotlight" of its attention moves across the audio, focusing on the relevant parts for each character it writes.

Conclusion

In this lesson, we have formally defined the problem of automatic speech recognition and framed it within the powerful sequence-to-sequence paradigm.

Key Takeaways:

  • ASR can be formulated as a sequence-to-sequence problem, mapping a long sequence of audio features () to a much shorter sequence of text tokens ().
  • The fundamental difficulty in ASR is the alignment challenge: the mapping between the continuous, variable-rate audio signal and the discrete text transcript is unknown.
  • Manually creating this alignment is impractical, so modern ASR models must learn it automatically.
  • Two dominant paradigms for automatic alignment are:
    • Connectionist Temporal Classification (CTC): Makes a prediction per audio frame and uses a blank token and a collapsing function to map the long prediction sequence to the short output sequence, marginalizing over all possible alignments.
    • Attention Mechanisms: An encoder-decoder model where the decoder learns to focus on the most relevant parts of the input audio at each step of generating the output.

Preview of the Next Lesson:

We have introduced the high-level concept of CTC as a solution to the alignment problem. In our next lesson, we will get into the details. We will derive the Connectionist Temporal Classification (CTC) loss function and understand the forward-backward algorithm that makes its computation tractable. This will give you a deep mathematical understanding of one of the most important algorithms in modern ASR.

Can't find a good explanation? Sign up and we'll make it for you

Sign up