Skip to main content
Create your own
Lesson illustration

STFT: Derivation and Windowing Trade-offs

Hello! Welcome to the fifth lesson in our module on Spectral Analysis.

In our last two lessons, we explored the Discrete Fourier Transform (DFT) and the highly efficient Fast Fourier Transform (FFT) algorithm used to compute it. We established that an FFT gives us a powerful snapshot of a signal's frequency content. However, this snapshot is static—it averages the frequencies over the entire duration of the signal. For a non-stationary signal like audio, where notes, words, and sounds are constantly changing, this is a significant limitation.

This lesson introduces the fundamental tool for analyzing time-varying frequency content: the Short-Time Fourier Transform (STFT). Our learning outcome is to derive the Short-Time Fourier Transform (STFT) and explain the trade-offs of different windowing functions. We will build the STFT from the ground up, starting with the DFT you already know. You'll learn how it creates a time-frequency representation of a signal and understand the critical role of window functions in controlling its behavior.

1. From a Still Image to a Movie of Frequencies

Audio signals are non-stationary; their spectral content is dynamic. A single DFT of an entire song tells you which notes were played on average, but not when they were played or in what sequence. We need to move from a static frequency "photograph" to a "movie" that shows how frequencies evolve over time.

The STFT achieves this through a simple and powerful strategy: divide and conquer. Instead of analyzing the whole signal at once, we break it into small, overlapping chunks (or "frames") and apply an FFT to each one. This assumes that over a very short duration (e.g., 20-40 milliseconds), the signal's properties are relatively stable, or "quasi-stationary."

Short-Time Fourier Transform Explained Easily

To start, let's get a high-level intuition for why the STFT is necessary and how it works. This video from Valerio Velardo's 'The Sound of AI' channel provides an excellent conceptual introduction.

Watch from the beginning to 06:30. The first part (to 03:13) explains the limitation of the DFT for audio. The second part (03:13 - 06:30) introduces the core STFT process of framing and the concept of a hop size.

This process can be visualized as a sliding window moving across the signal, with an FFT computed at each step.

Short-Time Fourier Transform (STFT) Process Illustration
This diagram illustrates the core STFT workflow. A long signal is segmented by overlapping windows. Each windowed segment is then fed into an FFT, producing a frequency spectrum for that specific moment in time.

2. Deriving the STFT Formula

Let's formalize this process mathematically. Recall the DFT formula for a signal of length :

To adapt this for the STFT, we need to incorporate the concepts of framing and windowing. We introduce two key modifications:

  1. A time index, , which denotes the current frame number.
  2. A window function, , which is a finite-length sequence that we multiply our signal frame by.

The STFT, denoted as , is a function of both the frame index and the frequency bin . It is calculated by taking the DFT of the product of the signal and a window function that is shifted to the -th frame.

The precise mathematical formulation can be a bit dense at first. The following video segment does an exceptional job of comparing the DFT formula side-by-side with the STFT formula, explaining each new component.

Short-Time Fourier Transform Explained Easily

Let's carefully examine the math. This segment from the same video breaks down the STFT formula, connecting it directly to the DFT formula we already know.

Watch from 09:34 to 17:16. Pay close attention to how the STFT formula introduces the frame index 'm' and how the summation is no longer over the entire signal but over a windowed frame. Understand what each part of the equation represents: the signal chunk, the window, and the complex exponential.

Based on the video, the discrete STFT is defined as:

Let's unpack this:

  • is the complex STFT value for frame and frequency bin .
  • is the input signal, shifted to the start of the -th frame. is the hop size (or hop length), the number of samples you slide the window forward for each new frame.
  • is the window function of length .
  • The summation is over the length of the window, .
  • is the size of the FFT used. This is often equal to the window length , but can be larger (by zero-padding) to get a higher-resolution frequency axis.

The output is a 2D matrix where one axis represents time (frame index ) and the other represents frequency (frequency bin ).

Short-Time Fourier Transform (STFT) Formula Breakdown
This image provides a clear, annotated version of the STFT formula, labeling the signal, window function, and the resulting time and frequency parameters.

3. The Crucial Role of Window Functions

The choice of the window function, , is not a minor detail; it is fundamental to the quality of your spectral analysis. The simplest choice is a rectangular window, which is 1 inside the frame and 0 everywhere else. However, the sharp transitions at the edges of this window introduce artifacts.

Multiplying a signal by a window in the time domain is equivalent to convolving their spectra in the frequency domain. The spectrum of a rectangular window is a sinc-like function, which has a narrow main lobe but fairly high "side lobes". This convolution causes energy from a single frequency to "leak" into adjacent frequency bins, an effect called spectral leakage.

To mitigate this, we use tapered windows that smoothly approach zero at their edges. This reduces the abruptness of the truncation, leading to much lower side lobes in the frequency domain.

Spectrum Analysis Windows

Let's dive deep into the world of window functions. This classic resource provides a comprehensive look at the properties and trade-offs of various windows. We'll start with the Rectangular window to understand the problem of spectral leakage.

Read the introduction and the section on 'The Rectangular Window'. Focus on understanding the terms 'main lobe' and 'side lobes' from the figures. Note the key properties: a main-lobe width of 4\pi/M and a first side-lobe level of only -13 dB.

The high side lobes of the rectangular window (-13 dB) mean that a strong frequency component can create leakage that is louder than a weak, nearby frequency component, completely obscuring it. This leads to a fundamental trade-off in window selection.

The Main-Lobe vs. Side-Lobe Trade-off

By choosing a smoother, more tapered window, you can dramatically reduce the height of the side lobes, but it comes at a cost: the main lobe gets wider.

  • Narrow Main Lobe: Better frequency resolution. You can distinguish between two sinusoidal components that are very close in frequency.
  • Low Side Lobes: Better dynamic range. You can detect a weak sinusoidal component in the presence of a strong one nearby.

Let's compare two of the most common tapered windows: Hann and Hamming.

Spectrum Analysis Windows

Now let's examine the solutions to the rectangular window's problems. This reading introduces the generalized Hamming family, including the popular Hann and Hamming windows.

Read the sections 'Generalized Hamming Window Family', 'Hann or Hanning or Raised Cosine', and 'Hamming Window'. Compare the key properties listed in the summaries: Main-lobe width First side-lobe level (in dB) Side-lobe roll-off rate (in dB/octave)

Here is a summary of the trade-offs:

Window Main-Lobe Width Highest Side Lobe Side-Lobe Roll-off Best For...
Rectangular Base -13 dB -6 dB/octave Strong signals with no close interferers.
Hann Base -31 dB -18 dB/octave General purpose, good balance.
Hamming Base -43 dB -6 dB/octave Excellent side-lobe rejection near main lobe.
Blackman Base -58 dB -18 dB/octave Applications needing very high dynamic range.

This trade-off is not just theoretical. The following reading shows a practical example of analyzing an oboe tone, where the choice of window directly impacts what you can "see" in the spectrum.

Spectrum Analysis Windows

Let's see this trade-off in action with a real audio signal.

Read the section 'Spectrum Analysis of an Oboe Tone'. Observe how the harmonics in the oboe spectrum, which are obscured by the side-lobes of the rectangular and Hamming windows, become clearly visible when using the Blackman window.

4. The Time-Frequency Uncertainty Principle

There is a second, equally important trade-off in STFT, controlled not by the window shape, but by the window size (or frame duration).

  • A short window gives you good time resolution. You can pinpoint when a frequency event occurred with high precision. However, it gives you poor frequency resolution because the analysis window is too short to capture enough cycles of a low-frequency wave.
  • A long window gives you good frequency resolution. By analyzing a longer segment, you can resolve closely spaced frequencies. However, this comes at the cost of poor time resolution, as any events that happen within that long window are blurred together in time.

This is a fundamental constraint of signal processing, often related to the Heisenberg Uncertainty Principle in physics. You cannot simultaneously have perfect resolution in both time and frequency.

[PDF] 3. Short-Time Fourier Transforms

The following document from Carnegie Mellon University provides a concise explanation and a powerful visual to illustrate this trade-off.

Please read the section '3.2.1 Impact of window size and shape'. Focus on the text explaining Figure 3.1. Observe how the short window (10 ms) produces vertical stripes (good time resolution, poor frequency resolution) and the long window (50 ms) produces horizontal stripes (good frequency resolution, poor time resolution).

5. An Alternative View: The STFT as a Filter Bank

Given your background, you may appreciate another perspective on the STFT. The entire operation is mathematically equivalent to passing the signal through a bank of parallel, overlapping bandpass filters.

  • Each filter is centered on a different frequency bin.
  • The shape of each filter's frequency response is determined by the Fourier transform of the window function, .
  • The output of the STFT at frame and frequency corresponds to the output of the -th filter at time .

From this viewpoint, a long window in time has a narrow Fourier transform, resulting in a bank of very narrow, selective filters (good frequency resolution). A short window has a wide Fourier transform, resulting in a bank of wide, less selective filters (good time resolution). This is just another way to understand the same time-frequency trade-off.

Short-time Fourier Transform and the Spectogram

This video from Barry Van Veen offers a clear explanation of this filter bank interpretation.

Watch the segment from 10:08 to 13:09. The visuals effectively show how the STFT can be modeled as a set of parallel bandpass filters and how the window length relates to the bandwidth of those filters.


Conclusion

In this lesson, we have moved from the static world of the DFT to the dynamic, time-varying analysis of the STFT. You've seen how it's constructed and have learned about the critical trade-offs involved in its application.

Key Takeaways:

  • The STFT overcomes the limitation of the DFT by analyzing short, overlapping frames of a signal to see how its frequency content changes over time.
  • The STFT formula is essentially a DFT applied to a windowed signal segment, with the result being a 2D matrix of time vs. frequency.
  • Window functions are applied to each frame to reduce spectral leakage. Tapered windows (like Hann, Hamming) are used to minimize this effect.
  • There is a fundamental trade-off in window shape: better frequency resolution (narrower main lobe) vs. better dynamic range (lower side lobes).
  • There is a fundamental trade-off in window size: better time resolution (short window) vs. better frequency resolution (long window).
  • The STFT can also be interpreted as a filter bank, where the window function's transform defines the shape of each bandpass filter.

In our next lesson, we will put this theory into practice. We will use Python libraries like torchaudio and librosa to compute and visualize spectrograms—the visual representation of the STFT's magnitude—which are the cornerstone of modern audio AI.

Can't find a good explanation? Sign up and we'll make it for you

Sign up