Skip to main content
Create your own
Lesson illustration

1D CNNs for Audio Feature Extraction

Hello! Welcome to the first lesson of our fourth module, "Sequence Modeling with Transformers."

In the previous module, we built a strong foundation in preparing, augmenting, and evaluating audio data. We concluded by learning how to measure the performance of an ASR model using Word Error Rate (WER) and Character Error Rate (CER), which are essential for tracking progress during training.

Now, we shift our focus from the data to the models themselves. Our goal is to build models that can process sequential data like audio. You've asked to understand these concepts from first principles, and that's exactly where we'll start. Instead of immediately using pre-computed features like mel-spectrograms (which we'll revisit later), we'll explore how a neural network can learn meaningful features directly from the raw audio waveform.

This lesson addresses the learning outcome: Explain how 1D Convolutional Neural Networks (CNNs) can act as feature extractors for raw audio waveforms. This is the first and a crucial step in many modern, high-performance audio models like Whisper and wav2vec 2.0.

1. The Challenge of Raw Audio

As we saw in our first module, a digital audio signal is a long one-dimensional array of numbers representing the amplitude of the sound wave at discrete time steps. A single second of audio sampled at 16kHz is a sequence of 16,000 numbers.

What if we tried to feed this directly into a standard Multi-Layer Perceptron (MLP), or a "fully connected" network?

To understand the shortcomings of this approach, let's watch a short segment from a lecture on 1D CNNs.

Lecture 3.2a: 1-Dimensional Convolutional Neural Networks: getting started

This video from the DLVU YouTube channel introduces the problem of applying standard deep learning models to sound waves and explains why a simple MLP is not a good fit.

Watch from the beginning to 03:18 and then from 04:42 to 06:28. Pay close attention to the reasons why an MLP is ill-suited for processing long sequences like audio.

As the video highlights, using a standard MLP for raw audio presents several major problems:

  • Massive Number of Parameters: A fully connected layer would require a weight for every input sample connected to every neuron. For a 1-second audio clip, this would mean millions of weights in the very first layer, making the model incredibly large and difficult to train.
  • Loss of Temporal Structure: An MLP treats the input as a flat vector. It has no inherent understanding that sample t comes just before sample t+1. The crucial ordering of the audio signal is lost.
  • Not Translation Invariant: If the network learns to recognize a sound (like the phoneme /a/) at the beginning of a clip, it would have to re-learn it from scratch to recognize it in the middle or at the end. The learned knowledge isn't portable across the time axis.

Clearly, we need an architecture that is more efficient and respects the temporal, sequential nature of audio. This is where convolutions come in.

2. Understanding the Convolution Operation

At its core, a convolution is a mathematical operation that combines two functions or sequences to produce a third. In our case, we combine the input signal (the audio waveform) with a small, learnable filter (also called a kernel) to produce a feature map.

The best way to build a deep intuition for this is visually.

But what is a convolution?

This classic video from 3Blue1Brown provides a superb visual explanation of the convolution operation. While it covers several applications, we will focus on the fundamental 'flipping and sliding' mechanism.

Please watch from the beginning until 08:33. Focus on these key ideas: How convolution is a way of 'mixing' or 'combining' two lists of numbers. The visual of flipping one list and sliding it across the other, calculating a sum of products at each step. The moving average example, which is a direct application of 1D convolution.

The key takeaway from the video is the "sliding window" process. The kernel, which is much smaller than the input, slides across the input sequence. At each position, it computes a dot product between its own values and the values of the input signal it's currently on top of. The result of this dot product becomes a single value in the output feature map.

This simple mechanism has profound implications for processing sequential data.

3. From Sliding Windows to Feature Extraction

So, how does this sliding window help us find meaningful patterns in audio? The magic lies in the fact that the values in the kernel are learnable weights. The network learns, through backpropagation, to set the kernel's weights to match specific patterns it needs to detect.

Let's make this concrete with an example.

Lecture 3.2a: 1-Dimensional Convolutional Neural Networks: getting started

Returning to the DLVU lecture, we'll now see a brilliant example of a hand-crafted 'silence neuron' that illustrates how a convolutional filter works as a pattern detector.

Watch the segment from 09:25 to 13:22, and then from 15:42 to 17:58. Focus on how the weights of the neuron (the filter) are designed to 'fire' (produce a high output) when the input pattern matches the filter's pattern.

This is the core idea of a 1D CNN as a feature extractor:

  • A filter (or kernel) is like a mini-pattern detector.
  • When the pattern in the input signal under the filter is similar to the pattern of the filter's weights, the dot product is high, causing the output neuron to "fire."
  • By sliding this filter across the entire audio signal (the convolution operation), we can detect that specific pattern anywhere it appears in time.

The output of this operation, the feature map, is a new sequence where high values indicate the presence and location of the feature that the filter was tuned to detect.

Wav2vec-like Architecture for Audio Feature Extraction
This diagram illustrates how a 1D CNN acts as an encoder. The raw audio waveform (X) is fed into a series of convolutional layers. Each layer extracts progressively more complex features, transforming the raw samples into a sequence of latent feature vectors (Z).

4. The Architecture of a 1D Convolutional Layer

Now let's formalize the structure of a 1D CNN and its key properties that make it so effective for audio.

Real-time implementation of a deep learning based Voice Activity Detector

This Master's thesis provides a concise, technical description of the components of a 1D convolutional layer. It's a great reference for understanding the specific parameters you'll encounter in frameworks like PyTorch.

Please read sections 2.2 'Forward step', 2.2.1 'Convolutions', and 2.2.2 'Activation Functions'. Pay special attention to the definitions of Filter size, Stride, and Zero padding.

Let's summarize the key architectural concepts and properties:

  • Local Connectivity: Each neuron in the output feature map is connected to only a small, local region of the input, defined by the kernel size. This contrasts with an MLP where every neuron is connected to every input. This is computationally efficient and reflects the nature of signals, where local patterns are fundamental.

  • Parameter Sharing: The exact same filter (the same set of weights) is used across all positions of the input sequence. This is the "sliding" part. It dramatically reduces the number of parameters and is the mechanism that enables the next, most crucial property.

  • Translation Equivariance: As the video mentioned, "if you do a translation in the input, you will get exactly the same translation in the output." This means if a sound pattern appears 200ms later in the audio, the feature indicating its presence will also appear 200ms later in the feature map. The model recognizes the pattern regardless of its position in time, solving a major weakness of MLPs.

A typical Conv1D layer in a deep learning framework is defined by several key hyperparameters discussed in the reading:

  • in_channels: The number of feature maps from the previous layer. For the very first layer processing mono raw audio, this is 1.
  • out_channels: The number of filters to learn. Each filter will produce its own feature map, so this determines the "depth" of the output.
  • kernel_size: The temporal width of the filter. A larger kernel looks at a wider time window at once.
  • stride: The step size the filter takes as it slides across the input. A stride greater than 1 results in downsampling, reducing the length of the output sequence.
  • padding: Adding zeros to the start and end of the input to control the output length, often to keep it the same as the input length.

After the convolution, a non-linear activation function (like ReLU) is applied to allow the network to learn more complex, non-linear relationships.

5. CNNs as the Foundation of Modern Audio Models

You might ask: are these 1D CNNs still relevant, or are they just a stepping stone to Transformers? The answer is that they are a critical component of most state-of-the-art architectures.

Models like Whisper and wav2vec 2.0 don't feed the raw audio directly to a Transformer. Instead, they first use a stack of 1D CNN layers as a feature encoder.

Speech Command Recognition with torchaudio

This short excerpt from a tutorial on speech command recognition makes an important practical point about the receptive field of the first convolutional layer.

Read the paragraph under 'Define the Network'. Note the connection between kernel size, sampling rate, and the receptive field in milliseconds.

The CNN's job is to take the very long raw audio sequence and transform it into a shorter, richer sequence of feature vectors.

  1. Input: Raw waveform (e.g., 160,000 samples for 10 seconds at 16kHz).
  2. CNN Feature Encoder: A stack of Conv1D layers processes this waveform. Through a combination of kernel operations and striding, it produces a sequence of high-dimensional vectors. For example, it might output one 768-dimensional vector for every 20ms of audio.
  3. Output: A sequence of feature vectors (e.g., 500 vectors of size 768 for the 10-second clip).

This sequence of learned features is then passed on to a Transformer, which excels at modeling the long-range contextual relationships between these features.

Audio Feature Extraction with CNNs for Transformer Models
This diagram shows the typical architecture of a modern speech model. The raw audio is first processed by a CNN Encoder, which extracts latent features. These features are then fed into a Transformer Encoder to model the broader context. The 1D CNN is the essential first step.

The key is that the CNN filters are not hand-designed. The entire network, including the CNN feature encoder and the subsequent Transformer, is trained end-to-end. The network learns the optimal filters for the task at hand, whether it's recognizing speech, identifying music genres, or detecting voice activity.

Conclusion

In this lesson, we've unpacked how 1D Convolutional Neural Networks can move beyond hand-crafted features and learn directly from raw audio waveforms. This is a fundamental concept in modern audio AI.

Key Takeaways:

  • Standard MLPs are unsuitable for raw audio due to their massive parameter count, lack of temporal awareness, and failure to be translation invariant.
  • A 1D CNN uses a small, sliding filter (kernel) to detect local patterns in the audio sequence.
  • The properties of local connectivity, parameter sharing, and translation equivariance make CNNs highly efficient and effective for sequential data like audio.
  • The filters' weights are learned via backpropagation, allowing the network to automatically discover the most useful features for a given task.
  • In state-of-the-art models, a stack of 1D CNNs often serves as a feature encoder, transforming the long raw waveform into a shorter, richer sequence of feature vectors for a subsequent model like a Transformer to process.

Preview of the Next Lesson:
The CNN is excellent at learning local features. But how do we model the relationships between these features over longer periods? Before jumping to Transformers, we'll first explore the classic architecture for this task: Recurrent Neural Networks (RNNs) and LSTMs. Understanding them will provide crucial context for why Transformers and their attention mechanism became so revolutionary.

Can't find a good explanation? Sign up and we'll make it for you

Sign up