Skip to main content
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
Audio AI Researcher & Developer
ยท
Module 6
Supervised Speech Recognition Models
1
Speech Recognition: Sequence-to-Sequence and Alignment Challenges
Formulate speech recognition as a sequence-to-sequence problem and explain the alignment challenge.
2
CTC Loss and Forward-Backward Algorithm
Derive the Connectionist Temporal Classification (CTC) loss function and its forward-backward algorithm.
3
CTC Decoding: Greedy & Beam Search
Implement greedy and beam-search decoding algorithms for a CTC output probability matrix.
4
Understanding LAS: An Attention-Based ASR Architecture
Describe the Listen-Attend-Spell (LAS) architecture as an example of an attention-based encoder-decoder ASR model.
5
Whisper Model: Architecture and Training
Describe the Whisper model architecture and its multitask, multilingual training strategy.
6
Fine-tuning Whisper for Custom Speech
Fine-tune a pretrained Whisper model on a custom speech dataset using the Hugging Face ecosystem.
7
Shallow Fusion for CTC-based Acoustic Models
Explain how a language model can be integrated with a CTC-based acoustic model using shallow fusion.
8
ASR System Trade-offs: CTC, Attention, and Hybrid Architectures
Compare the trade-offs between CTC, attention-based, and hybrid ASR systems.
Previous module
Foundations of Generative Modeling
Next module
Self-Supervised Speech Representation