Hello! Welcome to the sixth lesson in our module on Spectral Analysis of Audio Signals.
In our last lesson, we delved into the theory of the Short-Time Fourier Transform (STFT). We established that the STFT provides a "movie" of a signal's frequency content by analyzing short, overlapping frames. We also discussed the critical trade-offs involved in choosing window functions and window sizes.
Today, we transition from theory to practice. The STFT produces a 2D matrix of complex numbers, which isn't directly interpretable. Our goal is to transform this raw data into a powerful and intuitive visualization: the spectrogram. This is a fundamental skill, as spectrograms are the "images of sound" that feed into many modern audio AI models.
Our learning outcome is to compute and visualize spectrograms from audio signals using Python libraries like librosa or torchaudio. We will walk through the entire process, from loading an audio file to generating a publication-quality plot.

1. From STFT Output to a Visual Spectrogram
Recall from our previous lesson that the STFT, , gives us a complex number for each time frame and frequency bin . To create a meaningful visualization, we need to perform two important processing steps.
Step 1: Calculating the Magnitude
The complex number contains both magnitude and phase information. The magnitude, , represents the amplitude or "strength" of that frequency component at that point in time. While phase is crucial for reconstructing the signal, for many analysis and classification tasks, we are primarily interested in the magnitude. A plot of the STFT's magnitude over time and frequency is often called a magnitude spectrogram.
Step 2: Converting to the Decibel (dB) Scale
If we plot the raw magnitude, the result is often underwhelming. The vast majority of the plot might appear dark, with only a few bright spots. This is because the energy in audio signals is distributed over a very wide dynamic range, and our perception of loudness is logarithmic, not linear.
To address this, we convert the amplitude values to the decibel (dB) scale. This logarithmic transformation achieves two things:
- It compresses the dynamic range, making both quiet and loud components of the signal visible.
- It aligns the visualization more closely with human auditory perception.
The following reading provides a great overview of why these perceptual scales are so important in audio processing.
Audio Deep Learning Made Simple - Why Mel Spectrograms ...
Let's read about the perceptual basis for using logarithmic scales for both frequency and amplitude. This will help you understand why simply plotting the raw STFT output is not enough.
Read the sections 'How do humans hear frequencies?' and 'How do humans hear amplitudes?'. Focus on the core idea that human perception is logarithmic, which motivates the use of scales like Mel (for frequency, which we'll cover next lesson) and Decibels (for amplitude).
Now, let's see how to implement this process using two of the most important Python libraries in audio AI.
2. Spectrograms with librosa
librosa is a highly versatile and widely used Python package for music and audio analysis. It offers a rich set of tools for feature extraction and visualization.
The following video provides a clear, step-by-step walkthrough of generating a spectrogram with librosa. We will follow its process.
Audio Spectrogram In Python Using Librosa & Matplotlib | Audio Machine Learning For Beginners
Watch this video to see a complete example of creating a spectrogram using librosa and matplotlib.
Watch from 04:21 to 10:40. The video will guide you through the key functions: librosa.stft() to compute the Short-Time Fourier Transform. np.abs() to get the magnitude. librosa.amplitude_to_db() to convert to the decibel scale. librosa.display.specshow() to plot the final spectrogram.
Let's summarize the key steps in code. Assume you have an audio signal y and a sampling rate sr (e.g., from y, sr = librosa.load(audio_path)).
import librosa
import librosa.display
import numpy as np
import matplotlib.pyplot as plt
# 1. Define STFT parameters
# These parameters directly relate to our theory lesson.
n_fft = 2048 # FFT window size. Determines frequency resolution.
hop_length = 512 # Number of samples to slide the window. Determines time resolution.
# 2. Compute the STFT
# The output is a complex-valued matrix D
D = librosa.stft(y, n_fft=n_fft, hop_length=hop_length)
# 3. Separate magnitude and phase
# We typically visualize the magnitude
S_magnitude = np.abs(D)
# 4. Convert magnitude to decibels
S_db = librosa.amplitude_to_db(S_magnitude, ref=np.max)
# 5. Plot the spectrogram
fig, ax = plt.subplots(figsize=(10, 4))
img = librosa.display.specshow(S_db, sr=sr, hop_length=hop_length, x_axis='time', y_axis='log', ax=ax)
fig.colorbar(img, ax=ax, format='%+2.0f dB', label='Magnitude (dB)')
ax.set_title('Spectrogram (log-frequency)')
ax.set_xlabel('Time (s)')
ax.set_ylabel('Frequency (Hz)')
plt.show()
The librosa.display.specshow function is particularly powerful as it automatically handles the formatting of the time and frequency axes based on the sr and hop_length parameters. Using y_axis='log' provides a logarithmic frequency scale, which is often more intuitive for audio.
3. Spectrograms with torchaudio
For building deep learning models in PyTorch, torchaudio is the natural choice. It integrates seamlessly with PyTorch tensors and provides GPU-accelerated operations, which is essential for training. torchaudio approaches this task using a "transform" paradigm, which should be familiar from your experience with torchvision.
This next resource provides a complete guide to using torchaudio for this task.
Use TorchAudio to Prepare Audio Data for Deep Learning
This article from Real Python demonstrates the torchaudio workflow for creating and visualizing spectrograms. It's a great textual complement to the video-based librosa guide.
Read the sections 'Create Spectrograms' and 'Visualize Audio Spectrograms'. Pay attention to the object-oriented approach: you first create a transform object (e.g., T.Spectrogram), and then you call it on your data.
Here is the equivalent torchaudio code. Assume you have a waveform loaded as a PyTorch tensor, waveform, and its sample rate sample_rate.
import torch
import torchaudio
import torchaudio.transforms as T
import matplotlib.pyplot as plt
# Ensure waveform is a 2D tensor [channels, time]
if waveform.ndim == 1:
waveform = waveform.unsqueeze(0)
# 1. Define STFT parameters
n_fft = 2048
hop_length = 512
win_length = n_fft # The window length, often same as n_fft
# 2. Create the Spectrogram transform
# This object can be reused or placed in a nn.Sequential pipeline
spectrogram_transform = T.Spectrogram(
n_fft=n_fft,
win_length=win_length,
hop_length=hop_length,
normalized=True,
)
# 3. Apply the transform to get the magnitude spectrogram
S_magnitude = spectrogram_transform(waveform)
# 4. Create and apply the Amplitude-to-DB transform
# `stype='power'` or `stype='magnitude'` is important. Spectrogram gives magnitude.
# `top_db` is a clipping threshold to prevent -inf values.
db_transform = T.AmplitudeToDB(stype='magnitude', top_db=80.0)
S_db = db_transform(S_magnitude)
# 5. Plot the spectrogram
fig, ax = plt.subplots(figsize=(10, 4))
# S_db is [channels, freq_bins, time_frames], so we squeeze the channel dim
img = ax.imshow(S_db.squeeze(0).numpy(), aspect='auto', origin='lower',
extent=[0, waveform.shape[1] / sample_rate, 0, sample_rate / 2])
fig.colorbar(img, ax=ax, format='%+2.0f dB', label='Magnitude (dB)')
ax.set_title('Spectrogram (linear-frequency)')
ax.set_xlabel('Time (s)')
ax.set_ylabel('Frequency (Hz)')
plt.show()
Notice that with matplotlib.pyplot.imshow, you need to manually calculate the extent to label the axes correctly, unlike the convenience of librosa.display.specshow. Both libraries, however, will produce visually similar spectrograms if given the same parameters.
4. The Time-Frequency Trade-off in Practice
In the previous lesson, we discussed the time-frequency uncertainty principle: a short analysis window gives good time resolution but poor frequency resolution, while a long window does the opposite. Let's see this in action. The n_fft parameter (and win_length) controls the window size.
Intro to Audio Processing for Deep Learning
The following video segment provides an excellent practical demonstration of this trade-off by adjusting the window size and showing the effect on the resulting spectrogram.
Please watch from 29:39 to 39:09. The speaker experiments with different window sizes (nperseg in scipy, which is equivalent to n_fft) and clearly explains the visual results. Observe how the spectrogram changes from being 'smeared' horizontally (poor time resolution) to 'smeared' vertically (poor frequency resolution).
To summarize the visual effect:
- Large
n_fft(e.g., 4096): You are using a long window. The spectrogram will have very fine, detailed horizontal lines, clearly resolving individual frequency harmonics. However, fast temporal events (like a quick drum hit) will appear smeared out vertically across several time frames. Good frequency resolution, poor time resolution. - Small
n_fft(e.g., 256): You are using a short window. Fast temporal events will be sharply localized in time, appearing as crisp vertical lines. However, the frequency information will be blurry, with horizontal bands appearing thick and indistinct. Good time resolution, poor frequency resolution.
Choosing the right STFT parameters is a critical step in feature engineering for audio. The optimal choice depends on the specific characteristics of the sound you are analyzing. For speech, window sizes of 25-40 ms are common (n_fft = 400-640 at 16kHz). For music, where you might need to resolve musical notes precisely, longer windows are often used.
Conclusion
Congratulations! You have successfully bridged the gap between the theory of the STFT and its practical implementation. You can now generate, visualize, and interpret one of the most fundamental representations in audio processing.
Key Takeaways:
- A spectrogram is a visual representation of an audio signal's frequency content over time.
- It's computed by taking the magnitude of the STFT output and converting it to a decibel (dB) scale to match human perception and compress the dynamic range.
- Python libraries like
librosaandtorchaudioprovide robust tools to compute and visualize spectrograms.librosais excellent for general analysis, whiletorchaudiois optimized for PyTorch deep learning pipelines. - The choice of STFT parameters, especially the FFT window size (
n_fft), has a direct and visible impact on the time-frequency resolution trade-off of the resulting spectrogram.
In our next lesson, we will refine this representation even further. You'll learn about the mel scale, a psychoacoustically-inspired frequency scale that gives more importance to lower frequencies, just like the human ear. We will then learn to convert our linear spectrograms into mel-spectrograms, which are the de-facto standard input for a vast majority of speech recognition and audio classification models.