Skip to main content
Create your own

Activation Functions: A Comparative Study

Hello! Welcome to the third lesson in our module on Deep Neural Network Fundamentals.

Introduction

In our last session, you did a fantastic job deriving and implementing the backpropagation algorithm. We built a neural network that can learn by calculating the gradient of the loss function and updating its weights and biases. In our implementation, we used the ReLU activation for our hidden layer and Softmax for the output, choices that were effective but somewhat arbitrary at the time.

Today, we'll explore the "why" behind those choices. This lesson is dedicated to comparing the properties and use cases of the most common activation functions: Sigmoid, Tanh, ReLU, and Leaky ReLU.

Activation functions are a cornerstone of neural network design. They are the primary source of the non-linearity that allows deep networks to learn complex patterns. Your choice of activation function can dramatically affect training stability, convergence speed, and overall model performance. Understanding their trade-offs is crucial for designing and debugging deep learning models. This lesson will build directly on our backpropagation discussion, as many of the issues with activation functions reveal themselves through their gradients.

1. The Essential Role of Non-Linearity

As a quick refresher, why do we need activation functions at all? A neural network layer performs a linear operation: . If we were to stack these layers without any transformation in between, the entire network would still only be able to represent a single, complex linear function.

For a network to learn non-linear relationships—like classifying images or understanding language—we must introduce a non-linear function after each linear transformation. This is the role of the activation function, .

To get a quick visual intuition of why this is so critical, let's watch a short segment of the following video.

Activation Functions In Neural Networks Explained | Deep Learning Tutorial

This video from AssemblyAI provides a clear, concise explanation of why non-linearity is fundamental for learning complex patterns.

Please watch from 00:28 to 01:54. Focus on the conceptual difference between what a network can learn with and without activation functions.

2. The "Saturating" Duo: Sigmoid and Tanh

Historically, the first widely used non-linear activations were S-shaped functions that "squash" an infinite input range into a finite output range.

For a detailed breakdown of these functions, including their mathematical definitions, advantages, and limitations, the following article provides an excellent guide.

12 Types of Activation Functions in Neural Networks

This comprehensive guide from Medium discusses 12 different activation functions. We will use it as our primary reference for this lesson. Let's start with the sections on Sigmoid and Tanh.

Please read the sections titled '3. Sigmoid Function' and '4. Tanh (Hyperbolic Tangent) Function'. As you read, pay close attention to: The mathematical formulas and output ranges for both functions. The key improvement of Tanh over Sigmoid (being zero-centered). The major limitation they both share: the vanishing gradient problem.

Key Properties and Problems

Let's summarize and connect these points to our previous lesson on backpropagation.

Sigmoid (Logistic) Function

  • Formula:
  • Output Range: (0, 1). This makes it a natural choice for the output layer in binary classification, where the output can be interpreted as a probability.
  • Problem 1: Vanishing Gradients. The derivative of the sigmoid is . The maximum value of this derivative is 0.25 (at z=0). For large positive or negative inputs, the function saturates (flattens out), and its gradient approaches zero. During backpropagation, these small gradients are multiplied through many layers, causing the gradients for the initial layers to become infinitesimally small. This effectively stops them from learning.
  • Problem 2: Not Zero-Centered. The output is always positive. This means that during backpropagation, the gradients for the weights of a layer will all have the same sign (either all positive or all negative). This leads to inefficient, zig-zagging updates during gradient descent, slowing down convergence.

Tanh (Hyperbolic Tangent) Function

  • Formula: (which is just a scaled and shifted sigmoid).
  • Output Range: (-1, 1).
  • Improvement: Tanh is zero-centered. Its outputs can be positive or negative, which helps to mitigate the zig-zagging gradient updates seen with Sigmoid, generally leading to faster convergence.
  • Problem: It still suffers from the vanishing gradient problem. Like Sigmoid, it saturates at the extremes, and its gradients approach zero, making it difficult to train very deep networks.

For a deeper intuition on how the shape of the sigmoid's derivative leads to vanishing gradients, the following video segment is excellent.

Activation Functions - EXPLAINED!

The CodeEmporium channel offers a great visualization of why sigmoid's squeezing nature is problematic during backpropagation.

Watch from 07:08 to 08:29. This part clearly shows the derivative of the sigmoid function and explains how multiplying these small values layer by layer causes the gradient to 'vanish'.

Because of these issues, Sigmoid and Tanh are now rarely used in the hidden layers of modern deep feed-forward or convolutional networks.

3. The Modern Default: ReLU and its Variants

To overcome the vanishing gradient problem, a new family of activation functions emerged. The most popular among them is the Rectified Linear Unit, or ReLU.

Let's return to our reading to understand ReLU and its successor, Leaky ReLU.

12 Types of Activation Functions in Neural Networks

We'll continue with the same article to learn about ReLU and a key improvement, Leaky ReLU.

Please read the sections titled '5. ReLU (Rectified Linear Unit)' and '6. Leaky ReLU'. Focus on: ReLU's simple formula and its advantages (no vanishing gradient for positive inputs, computational efficiency). ReLU's primary limitation: the 'Dying ReLU' problem. How Leaky ReLU's simple modification addresses this problem.

Key Properties and Problems

ReLU (Rectified Linear Unit)

  • Formula:
  • Advantages:
    1. No Vanishing Gradient (for positive inputs): For , the derivative is a constant 1. This means the gradient can be passed back through many layers without shrinking, which greatly accelerates the training of deep networks.
    2. Computational Efficiency: It's a very simple operation (a thresholding at zero), which is much faster to compute than the exponentials in Sigmoid and Tanh. This is a significant advantage given your background in software engineering, as you know that seemingly small per-operation costs can add up massively at scale.
    3. Sparsity: For any given input, many neurons will have a negative pre-activation and thus an output of 0. This "sparse activation" can make the network more efficient and can lead to more robust, disentangled representations.
  • Problem: The "Dying ReLU" Problem. For , the derivative is 0. If a neuron's weights get updated in such a way that its pre-activation is consistently negative for all inputs in a batch, that neuron will have a gradient of 0. It will stop updating its weights and effectively "die," playing no further part in the network's learning.

Leaky ReLU

  • Formula: , where is a small constant (e.g., 0.01).
  • Improvement: It solves the dying ReLU problem. By introducing a small, non-zero gradient () for negative inputs, it ensures that neurons can always receive a gradient and continue to learn, even if their output is negative.

In practice, Leaky ReLU is often a safer choice than standard ReLU, although ReLU remains extremely popular and often works just as well.

Test your understanding!

You are training a very deep 100-layer neural network for an image classification task. During training, you notice that the loss stops decreasing after a few epochs, and the accuracy is stuck at a very low value. When you inspect the activations of the hidden layers, you find that in many layers, over 95% of the neurons are outputting zero for all images in your validation set.

  1. What activation function are you likely using in your hidden layers?
  2. What specific problem are you observing?
  3. What is a simple change you could make to your activation function to likely mitigate this issue?
Show answer
  1. You are likely using the ReLU activation function.
  2. The problem is the "Dying ReLU" problem. A large number of neurons have become "stuck" in the negative region, where their gradient is always zero, so they can no longer learn or update their weights.
  3. A simple and effective change would be to switch from ReLU to Leaky ReLU. This would ensure a small, non-zero gradient for all neurons, preventing them from dying and allowing them to recover and continue learning.

4. Summary and Practical Guidance

Choosing the right activation function is a key part of model design. Let's consolidate what we've learned.

The "Neural Network Activation Function Cheat Sheet" below provides a fantastic summary of the functions we've discussed.

Neural Network Activation Function Cheat Sheet
A concise cheat sheet comparing the properties, problems, and use cases of common activation functions. Note that Tanh, like Sigmoid, also suffers from vanishing gradients.

Rules of Thumb:

  • Hidden Layers: Start with ReLU or Leaky ReLU. They are the modern standard for most deep learning tasks, especially in computer vision (CNNs). They are computationally efficient and help mitigate the vanishing gradient problem.
  • Output Layer: The choice depends on your task.
    • Binary Classification: Use Sigmoid to output a probability between 0 and 1.
    • Multi-class Classification: Use Softmax (as we did in our implementation) to output a probability distribution across all classes.
    • Regression (predicting a continuous value): Use no activation (a linear function) so the output is not constrained to a specific range.
  • When to use Sigmoid/Tanh in hidden layers? Almost never in modern deep feed-forward or convolutional networks. Their primary modern use is in certain Recurrent Neural Network (RNN) architectures (which we will cover later), where their ability to "gate" information in a bounded range is useful.

5. Optional: A Glimpse at the State-of-the-Art

The quest for better activation functions is an active area of research. While ReLU and its variants are excellent general-purpose choices, several more advanced functions have been developed that often provide superior performance, especially in state-of-the-art models.

12 Types of Activation Functions in Neural Networks

For those interested in a deeper dive, let's explore some of these advanced functions. You will encounter these in modern papers and architectures, especially Transformers like BERT and GPT.

This is optional, but I highly recommend skimming the sections on: '7. Parametric ReLU (PReLU)': An extension of Leaky ReLU where the slope for negative inputs is learned during training. '8. Exponential Linear Unit (ELU)': An alternative to Leaky ReLU with a smooth exponential curve instead of a sharp corner. '10. Gaussian Error Linear Unit (GELU)': A smooth, non-monotonic function that has become the standard in modern Transformer models.

These functions—PReLU, ELU, SELU, GELU, Swish—each offer different trade-offs in terms of smoothness, computational cost, and performance, but they all build on the same core ideas we've discussed today: preventing dead neurons and ensuring healthy gradient flow.

Conclusion

Great work today! You now have a solid, principled understanding of the most important activation functions and the crucial role they play in deep learning.

Key Takeaways:

  • Activation functions introduce the non-linearity necessary for learning complex patterns.
  • "Saturating" functions like Sigmoid and Tanh were early choices but suffer from the vanishing gradient problem, making them unsuitable for deep hidden layers.
  • ReLU became the default due to its computational efficiency and its solution to the vanishing gradient problem for positive inputs.
  • ReLU's main drawback is the "Dying ReLU" problem, where neurons can get stuck and stop learning.
  • Leaky ReLU is a simple and effective fix for the dying ReLU problem and is an excellent all-around choice for hidden layers.
  • The choice of activation function for the output layer is determined by the specific problem you are solving (e.g., Sigmoid for binary classification, Softmax for multi-class).

Preview of the next lesson:
Our choice of activation function has a direct impact on another critical aspect of training: weight initialization. A network with random weights is in a fragile state. If the initial weights are too large or too small, the signals passing through the network can rapidly shrink to zero or explode to infinity, especially when using activations like ReLU. In the next lesson, we will explore initialization strategies like Xavier and He initialization, which are specifically designed to work with particular activation functions to ensure a healthy gradient flow right from the start of training.

Can't find a good explanation? Sign up and we'll make it for you

Sign up