Skip to main content
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
Audio AI Researcher & Developer
ยท
Module 11
Model Optimization for Deployment
1
Knowledge Distillation for Speech Model Compression
Explain the principles of knowledge distillation and apply it to compress a large speech model into a smaller student model.
2
Quantizing Speech Models: Dynamic and Static Approaches
Apply post-training dynamic quantization and static quantization to a speech model.
3
Quantization Trade-offs: Size, Speed, and Performance
Evaluate the trade-off between model size, inference speed, and performance degradation after quantization.
4
PyTorch to ONNX for Speech Models
Export a trained PyTorch speech model to ONNX format for framework-agnostic deployment.
5
Model Inference and Performance Benchmarking with ONNX Runtime
Perform inference with the exported model using ONNX Runtime and benchmark its performance.
6
Streaming ASR Architecture and Latency-Accuracy Trade-offs
Describe the architectural requirements for streaming ASR systems and their latency-accuracy trade-offs.
7
Streaming ASR Inference Loop
Implement a basic streaming inference loop for a frame-synchronous ASR model.
8
Optimizing Speech AI Inference: Identifying Bottlenecks
Analyze the performance bottlenecks in a typical speech AI inference pipeline.
Previous module
Audio Language Models and S2S Translation
Next module
Building and Deploying Speech Services