Hello! Welcome to your next lesson in our journey through audio AI.
In the previous lesson, we fine-tuned Whisper, an encoder-decoder model that benefits from a powerful language model learned implicitly during its massive pre-training. This architecture, where the language model is an integral part of the trained network, is often called deep fusion.
However, many other state-of-the-art Automatic Speech Recognition (ASR) systems, especially those based on Connectionist Temporal Classification (CTC) like wav2vec 2.0, are trained primarily as acoustic models. They excel at mapping audio to a sequence of characters but lack explicit, high-level linguistic knowledge. This can lead to errors that are phonetically plausible but linguistically nonsensical (e.g., transcribing "I read a book" as "I red a book").
Today's lesson addresses this gap by focusing on our learning outcome: Explain how a language model can be integrated with a CTC-based acoustic model using shallow fusion. We will explore the theory, mathematics, and practical implementation of this widely-used technique to improve ASR accuracy.
1. The Need for a Language Model
An acoustic model (AM) answers the question, "What does this audio sound like?" It computes the probability of a character sequence given an audio input, . A language model (LM) answers a different question: "Is this a probable sequence of words in this language?" It computes the probability of a sequence of words, .
By combining these two, we can correct errors where the acoustics are ambiguous but the linguistic context is clear.
I Built a Personal Speech Recognition System for my AI Assistant
To see a clear, practical example of why a language model is so important, watch this segment from the video "I Built a Personal Speech Recognition System for my AI Assistant" by The AI Hacker.
Please watch from 04:04 to 06:38 and then from 10:51 to 12:11. The first part explains the conceptual role of the LM and introduces CTC beam search as the integration mechanism. The second part provides a powerful before-and-after demo, showing the dramatic improvement in transcription quality when an LM is used.
As the video demonstrates, the acoustic model alone can easily make mistakes with homophones ("read" vs. "red") or produce grammatically awkward sentences. The language model acts as a "sense checker," rescoring the possible transcriptions to favor those that are linguistically more plausible.
The mechanism that enables this combination during inference is beam search decoding.
2. The Role of Beam Search Decoding
In our lesson on CTC, you learned about greedy decoding, which simply takes the most probable character at each time step. This is fast but suboptimal because it can't recover from early mistakes.
Beam search is a more effective decoding algorithm. Instead of keeping only the single most likely sequence, it maintains a "beam" of the top most probable hypotheses (or "beams") at each step.
End-to-End Speech Model Design
This blog post by Arun Baby provides a concise explanation of beam search and how it overcomes the limitations of greedy decoding.
Read the section titled 'Deep Dive: Beam Search Decoding'. Pay attention to the example of how it avoids getting stuck with 'The read apple' by keeping multiple hypotheses alive.
At each step of the beam search, for each of the hypotheses currently in the beam, the algorithm proposes new hypotheses by extending them with possible next characters. It then scores these new, longer hypotheses and keeps only the top overall. It's during this scoring step that we can integrate the language model.
3. The Mathematics of Shallow Fusion
Shallow fusion is an inference-time technique that combines the scores from the acoustic model and an external language model in a linear interpolation. The core idea is to find the text sequence that maximizes a combined score.
Since probabilities are often very small, it's standard practice to work with log-probabilities to maintain numerical stability. The combined score for a hypothesis given an audio input is calculated as follows:
Let's break this down.

-
: This is the acoustic model score. It's the log-probability of the hypothesis given the audio , as calculated by the CTC model. This score comes directly from the neural network's outputs.
-
: This is the language model score. It's the log-probability of the hypothesis according to an external, pre-trained LM (like KenLM). This score reflects how grammatically and semantically likely the sentence is.
-
(lambda, often called
lm_weight): This is the language model weight. It's a crucial hyperparameter that balances the influence of the AM and the LM.- If is too low, the LM's contribution is negligible.
- If is too high, the system may output grammatically perfect but acoustically incorrect sentences, ignoring what was actually said.
- Typical values are tuned on a validation set and often fall in the range of 0.5 to 2.0.
-
: This is the word insertion bonus. CTC models have an inherent bias towards shorter sequences because they can merge repeated characters. This term adds a small bonus for each word in the hypothesis, counteracting the bias and preventing the decoder from favoring overly short (and often incorrect) transcriptions. is another hyperparameter to be tuned.
This entire calculation happens within the beam search algorithm at each step to score and prune candidate hypotheses.

4. Practical Implementation with torchaudio
Now let's see how this theoretical concept is put into practice. The torchaudio library provides a powerful and efficient CTC beam search decoder that has built-in support for shallow fusion with a KenLM language model.
ASR Inference with CTC Decoder — Torchaudio 2.6.0 documentation
The PyTorch documentation offers an excellent tutorial on using the CTC decoder. We'll walk through the key parts to see how the components we've discussed are used in code.
Read the following sections from this tutorial: 'Overview', 'Language Model', 'Construct Decoders', 'Run Inference', and 'language model weight'. As you read, focus on: The four components required: Acoustic Model, Tokens, Lexicon, and Language Model. How the ctc_decoder function is initialized with an LM file. The clear improvement in the output when using the beam search decoder with an LM versus the greedy decoder. The explanation of the lm_weight parameter, which directly corresponds to the \lambda in our formula.
Let's summarize the key steps from the torchaudio tutorial:
-
Gather Components: You need four things:
- Acoustic Model: A trained CTC model, like
wav2vec2.0-base-960h. - Tokens: A list of the possible characters the model can predict.
- Lexicon: A mapping of words to their character sequences (e.g.,
HELLO H E L L O). This constrains the search to valid words. - Language Model: A KenLM n-gram model file (e.g.,
lm.bin). N-gram models are statistical (not neural), which makes them very fast and memory-efficient for inference.
- Acoustic Model: A trained CTC model, like
-
Construct the Decoder: You instantiate the decoder using the
torchaudio.models.decoder.ctc_decoderfactory function, passing the paths to your tokens, lexicon, and LM files, along with the hyperparameters we discussed.from torchaudio.models.decoder import ctc_decoder beam_search_decoder = ctc_decoder( lexicon="path/to/lexicon.txt", tokens="path/to/tokens.txt", lm="path/to/lm.bin", # The Language Model nbest=1, beam_size=1500, lm_weight=1.8, # This is our lambda (λ) word_score=-1.0, # This contributes to our beta (β) ) -
Run Inference: You first get the emissions (log-probabilities) from your acoustic model for a given audio waveform. Then, you pass these emissions to the decoder object.
# Assume `model` is your wav2vec 2.0 model and `waveform` is your audio with torch.no_grad(): emissions, _ = model(waveform) # The decoder performs the shallow fusion beam search hypotheses = beam_search_decoder(emissions) # Print the best hypothesis print(hypotheses[0][0].words)
As the tutorial demonstrates, the result from the beam search decoder with an LM is significantly more accurate and linguistically coherent than the output from a simple greedy decoder.
5. Shallow Fusion in Context
Shallow fusion is an effective and popular method, but it's not the only way to integrate an LM.
End-to-End Speech Model Design
Let's revisit the blog post from Arun Baby to see where shallow fusion fits in the broader landscape of LM integration techniques.
Read the section 'Deep Dive: Integrating Language Models'. It provides a succinct comparison of Shallow Fusion, Deep Fusion, and Cold Fusion.
To summarize the trade-offs:
- Shallow Fusion:
- Pros: Simple, flexible (you can swap LMs without retraining the AM), computationally cheap at training time (since it's an inference-only technique).
- Cons: The AM is not aware of the LM during training, so the two models are not jointly optimized.
- Deep Fusion:
- Pros: Potentially higher accuracy as the AM and LM are more tightly integrated during training.
- Cons: More complex to implement, requires retraining the entire system, less flexible.
For many applications, the simplicity and effectiveness of shallow fusion make it an excellent choice.
Conclusion
In this lesson, we have demystified the process of integrating an external language model with a CTC-based acoustic model. You now understand not just why this is necessary but also how it works, from the high-level concept down to the mathematical formulation and practical code implementation.
Key Takeaways:
- Acoustic Models (AMs) predict what an audio signal sounds like, while Language Models (LMs) predict what sentences are linguistically probable.
- Shallow fusion combines the scores from a CTC-based AM and an external LM during beam search decoding to produce more accurate transcriptions.
- The core formula involves a weighted sum of log-probabilities: .
- Hyperparameters like the language model weight () and a word insertion bonus () are critical for balancing the two models' influence.
- Libraries like
torchaudioprovide ready-to-use implementations of CTC beam search decoders with shallow fusion capabilities.
Preview of the Next Lesson:
We have now explored two dominant paradigms in ASR: the end-to-end encoder-decoder approach (Whisper) and the CTC-based approach augmented with shallow fusion (wav2vec 2.0). In our next lesson, we will directly "Compare the trade-offs between CTC, attention-based, and hybrid ASR systems." This will synthesize your knowledge and give you a framework for choosing the right architecture for a given ASR task.