Skip to main content
Create your own

Dilated Convolutions and Squeeze-and-Excitation Networks

Hello! Welcome to the next lesson on building powerful Convolutional Neural Networks.

Introduction

In our last lesson, we focused on efficiency, deconstructing the standard convolution into the depthwise separable convolution used in MobileNet. The main goal was to drastically reduce computational cost and model size for resource-constrained environments.

Today, we shift our focus from just efficiency to enhancing the representational power of our networks. We'll explore two advanced techniques that allow CNNs to extract more meaningful and context-aware features. Your learning outcome for this lesson is to apply dilated convolutions and squeeze-and-excitation networks for advanced feature extraction.

We will cover:

  1. Dilated (or Atrous) Convolutions: A method to increase a layer's receptive field—what it can "see"—without increasing computational cost or losing spatial resolution.
  2. Squeeze-and-Excitation (SE) Networks: A clever "plug-and-play" module that allows a network to perform channel-wise attention, learning to emphasize important feature channels and suppress less useful ones.

While the previous lesson was about making convolutions cheaper, this one is about making them smarter—first by expanding their spatial awareness, and then by refining their channel-wise focus.

1. Expanding the View: Dilated Convolutions

In many computer vision tasks, like semantic segmentation, we need to make a prediction for every single pixel in an image. However, standard CNNs often use pooling layers or strided convolutions to reduce the size of feature maps, which helps in classification but results in a loss of detailed spatial information.

How can we give our convolutional filters a larger receptive field to capture broader context, without downsampling the feature map or adding a huge number of parameters? The answer is dilated convolution, also known as atrous convolution.

The core idea is to introduce "holes" into a standard convolution kernel. A new parameter, the dilation rate (r), controls the spacing between kernel weights.

  • A rate of r=1 is a standard convolution.
  • A rate of r=2 means the kernel weights are applied to pixels that are one-step apart, effectively skipping every other pixel.

To understand this concept visually and mathematically, the following resource provides an excellent explanation.

DeepLabv3 & DeepLabv3+ The Ultimate PyTorch Guide

The article 'DeepLabv3 & DeepLabv3+ The Ultimate PyTorch Guide' from LearnOpenCV explains the mechanics of atrous (dilated) convolution clearly.

Please read the section titled 'Atrous Convolution'. Pay close attention to: The diagram visualizing a 3x3 kernel with a dilation rate of 2. The animation showing how the filter 'jumps' over the input feature map. The mathematical formula defining the operation.

As you can see, a 3x3 kernel with a dilation rate of r has the same number of parameters as a standard 3x3 kernel, but its receptive field is equivalent to a (2r+1) x (2r+1) kernel. This allows us to aggregate context from a wider area while keeping the computation and parameter count low, and importantly, without reducing the output resolution.

Test your understanding!

A standard 3x3 convolutional kernel has a receptive field of 3x3. What is the size of the receptive field of a 3x3 kernel with a dilation rate of r=4?

Show answer

The effective size of the receptive field can be calculated as k' = k + (k-1)(r-1), where k is the kernel size and r is the dilation rate.

For a 3x3 kernel (k=3) with r=4:
k' = 3 + (3-1)(4-1) = 3 + 2 * 3 = 9.
So, the receptive field is 9x9. It has only 9 parameters but covers the same area as a dense 9x9 kernel (which would have 81 parameters).

Application: Atrous Spatial Pyramid Pooling (ASPP)

The real power of dilated convolutions is unlocked when we use multiple dilation rates in parallel. This is the idea behind Atrous Spatial Pyramid Pooling (ASPP), a key component of the DeepLab family of segmentation models.

ASPP applies several parallel convolutions with different dilation rates to the same input feature map and then concatenates their results. This allows the model to probe the incoming features with filters of multiple receptive field sizes simultaneously, capturing objects and image context at various scales.

The following resource explains this module.

DeepLabv3 & DeepLabv3+ The Ultimate PyTorch Guide

Let's return to the LearnOpenCV article to see how dilated convolutions are used in practice within an ASPP module.

Read the section 'Atrous Spatial Pyramid Pooling (ASPP)'. Focus on: The diagram showing parallel filters with different rates capturing multi-scale features. The problem identified with very large dilation rates. The final modified ASPP architecture used in DeepLabv3, which includes a 1x1 convolution and an image-level feature branch to capture both local and global context.

By using ASPP, a network can effectively understand both fine-grained details and the broader scene context, which is critical for tasks like semantic segmentation where you need to distinguish a car from the road it's on.

2. Refining the Focus: Squeeze-and-Excitation Networks

We've just seen how to expand a filter's spatial view. Now, let's turn our attention inward, to the channels of a feature map. A convolution layer produces many channels, each representing a different learned feature. But are all these features equally important for the task at hand?

Probably not. Squeeze-and-Excitation (SE) Networks introduce a mechanism for the network to learn this importance dynamically. An SE block acts as a channel-wise attention module that adaptively recalibrates channel feature responses by explicitly modeling the interdependencies between channels. In simple terms, it learns a weight for each channel, boosting the important ones and diminishing the irrelevant ones.

The SE Block: Squeeze, Excite, Scale

The SE block is a lightweight, "plug-and-play" unit that can be added to existing CNN architectures like ResNet or Inception. It performs its function in three steps.

Squeeze-and-Excitation Block Diagram
A detailed block diagram of the Squeeze-and-Excitation (SE) block. It shows the three stages: Squeeze (Global Average Pooling), Excitation (a two-layer MLP bottleneck), and Scale (element-wise multiplication with the original input).

Channel Attention and Squeeze-and-Excitation Networks

The article 'Channel Attention and Squeeze-and-Excitation Networks' provides a fantastic, in-depth look at the SE block.

Please read the introduction and the 'Squeeze-and-Excitation Networks' section. As you read, focus on understanding the purpose of each of the three modules: Squeeze Module: Why is Global Average Pooling (GAP) used to shrink each channel's spatial dimensions (H x W) to a single number? Excitation Module: How does the two-layer MLP 'bottleneck' (with a reduction ratio r) learn the relationships between channels and generate attention weights? Scale Module: How are the learned attention weights (passed through a sigmoid function) applied to the original feature maps to perform the re-weighting?

In essence, the SE block computes a "signature" for each channel (Squeeze), determines a set of channel-wise weights based on these signatures (Excite), and then applies these weights to the original feature maps (Scale). This whole process is differentiable and trained end-to-end with the rest of the network.

Implementation and Integration

Given your background in software engineering, seeing how this abstract block translates to code will be very insightful. The following video provides a clear walkthrough of an implementation in TensorFlow.

Squeeze and Excitation Network Implementation in TensorFlow | Channel-wise Attention Mechanism

The video 'Squeeze and Excitation Network Implementation in TensorFlow' by Idiot Developer shows how to build an SE block from scratch.

Watch the implementation part of the video, from 5:34 to 11:26. Observe how the theoretical concepts map directly to TensorFlow/Keras layers: Squeeze: Implemented with GlobalAveragePooling2D. Excite: Implemented with two Dense layers, with the first one reducing the number of channels (the bottleneck) and the second one restoring it, followed by a sigmoid activation. Scale: A simple element-wise multiplication.

The beauty of the SE block is its modularity. It can be seamlessly integrated into many existing state-of-the-art architectures.

Schema of Inception vs. SE-Inception Module
A schematic showing how a Squeeze-and-Excitation block is added to a standard Inception module. The SE block operates on the concatenated output of the parallel branches before it's passed to the next layer.

The video below briefly demonstrates how to incorporate SE blocks and discusses their impact on performance.

Squeeze-and-Excitation | Lecture 11 | Applied Deep Learning

Let's watch a segment from Maziar Raissi's lecture on 'Squeeze-and-Excitation'.

Please watch from 6:05 to 8:30. This part covers two key points: How to integrate SE blocks into architectures like Inception and ResNet. The significant performance improvements (lower error rates) that SE blocks provide across a range of models, with very little additional computational cost (FLOPs).

The results are compelling: adding SE blocks consistently improves model accuracy on tasks like image classification and object detection, demonstrating the power of learning to focus on the most informative features.

Conclusion

In this lesson, we've added two powerful tools to our CNN toolkit that go beyond the standard convolution to enhance feature extraction.

Key Takeaways:

  • Dilated (Atrous) Convolution increases the receptive field without downsampling or adding significant cost. It's defined by a rate parameter that introduces gaps in the kernel.
  • Atrous Spatial Pyramid Pooling (ASPP) is a key application of dilated convolution, using parallel filters with different rates to capture multi-scale context, crucial for dense prediction tasks like semantic segmentation.
  • Squeeze-and-Excitation (SE) Networks provide a form of channel attention, allowing a model to adaptively re-weight feature channels based on their learned importance.
  • The SE Block is a lightweight, plug-and-play module that works in three steps: Squeeze (global information embedding via GAP), Excite (learning channel relationships with an MLP bottleneck), and Scale (re-weighting the input feature maps).

Preview of the next lesson:
This lesson marks the end of our module on the fundamental building blocks of CNNs. We've journeyed from classic architectures like LeNet and VGG, through the innovations of ResNet and Inception, to the efficiency of MobileNet, and finally to the advanced feature extraction techniques of today.

In the next module, we will begin our exploration of "Advanced Computer Vision Applications." We'll see how these powerful blocks are assembled into sophisticated systems to solve complex problems. We'll start by tackling object detection, implementing and analyzing region-based approaches like R-CNN and its faster successors.

Can't find a good explanation? Sign up and we'll make it for you

Sign up