Skip to main content
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
Audio AI Researcher & Developer
ยท
Module 3
Audio Data Augmentation and Pipelines
1
Speech Denoising with Spectral Gating and Wiener Filtering
Apply spectral gating and statistical Wiener filtering to denoise speech signals.
2
Voice Activity Detection (VAD) for Audio Segmentation
Perform voice activity detection (VAD) to segment speech from non-speech regions in an audio stream.
3
Time-Domain Audio Augmentation
Implement time-domain audio augmentation techniques, including speed perturbation and pitch shifting.
4
SpecAugment for Audio Data Augmentation
Implement frequency-domain audio augmentation by applying SpecAugment to mel-spectrograms.
5
Robust Audio Data Pipelines with torchaudio
Design and build a robust audio data loading and batching pipeline in PyTorch using torchaudio.
6
On-the-Fly Feature Extraction in PyTorch Data Pipelines
Integrate on-the-fly feature extraction (e.g., mel-spectrogram generation) into a PyTorch data pipeline.
7
Mastering Forced Alignment: From Theory to Timestamps
Explain the importance of forced alignment and apply a pre-trained model to generate word-level timestamps.
8
ASR Evaluation: WER and CER
Implement and interpret standard ASR evaluation metrics, specifically Word Error Rate (WER) and Character Error Rate (CER).
Previous module
Spectral Analysis of Audio Signals
Next module
Sequence Modeling with Transformers