Skip to main content
Create your own

Building an MLP: Forward Propagation

Hello! Welcome to the first lesson in our new module, "Deep Neural Network Fundamentals."

Introduction

In our last module, we explored unsupervised learning, culminating in the implementation of Gaussian Mixture Models with the Expectation-Maximization algorithm. We saw how EM can find the parameters of a probabilistic model with latent variables. Now, we pivot to the heart of modern AI: deep learning and supervised learning.

Today, we'll build our very first neural network: the Multi-Layer Perceptron (MLP). Specifically, we'll focus on forward propagation, which is the process a network uses to turn an input (like an image) into a prediction. Think of it as the network's "thinking" phase. We'll leave the "learning" phase, which involves adjusting the network's parameters, for our next lesson on backpropagation.

By the end of this lesson, you will be able to build a multi-layer perceptron from scratch and implement the forward propagation algorithm to generate predictions.

1. The Building Blocks: Neurons and Layers

At its core, a neural network is composed of simple computational units called neurons, organized into layers. Let's break down how they work.

The Single Neuron

Inspired by biological neurons, an artificial neuron takes several inputs, computes a value, and produces an output. This process has two main steps:

  1. Linear Combination: The neuron calculates a weighted sum of its inputs and adds a "bias" term. Given inputs and corresponding weights , this step computes a value , often called the pre-activation or logit.

    The weights () determine the importance of each input, while the bias () acts as an offset, allowing the neuron to activate even with zero inputs.

  2. Activation Function: The value is then passed through a non-linear activation function to produce the neuron's final output, or "activation," .

This non-linear step is crucial. Without it, stacking multiple layers of neurons would be mathematically equivalent to a single linear transformation. The non-linearity is what allows MLPs to learn and approximate highly complex, non-linear relationships in data.

To get a clear picture of these fundamentals, let's turn to a helpful guide.

Forward Propagation in Neural Networks: A Complete Guide

The article 'Forward Propagation in Neural Networks: A Complete Guide' from DataCamp provides an excellent breakdown of these foundational concepts.

Please read the section 'Foundations of Forward Propagation'. Pay close attention to: The biological inspiration and its computational analog. The mathematical formula for a single neuron's output. The role of the activation function in introducing non-linearity.

From Neurons to Layers

An MLP organizes these neurons into layers:

  • An input layer, which receives the raw data.
  • One or more hidden layers, which perform most of the computation and feature extraction.
  • An output layer, which produces the final prediction.

When we have a layer with multiple neurons, we can compute all their activations simultaneously using matrix operations—a process well-suited for your background in computer science and programming.

If we have an input vector (which is the activation from the previous layer, ), a weight matrix and a bias vector for layer , the computations for the entire layer are:

  1. Pre-activation:
  2. Activation:

Here, is the activation function for layer . This matrix formulation is not just elegant; it's the key to the efficient implementations found in all modern deep learning frameworks, leveraging the power of GPUs.

2. Forward Propagation: The Full Picture

Forward propagation is simply the process of chaining these layer-by-layer calculations from the input layer all the way to the output layer. The activations of one layer become the inputs for the next.

Forward Propagation in a Simple Neural Network
This diagram shows the flow of information in forward propagation. Starting with the input, each layer computes a weighted sum and applies an activation function. The process continues until the output layer produces a prediction, which is then compared to the true label to calculate the loss.

The general algorithm for a network with layers is:

  1. Start with the input data, .
  2. For each layer from 1 to :
    • Compute the pre-activation:
    • Compute the activation:
  3. The final activation, , is the network's prediction, often denoted as .

This process can be applied to a single data sample or, more commonly, to a "batch" of multiple samples at once. The vectorized approach is key for efficiency.

Forward Propagation in a Deep Network (C1W4L02)

To see a formal and clear explanation of this process, both for a single example and for a vectorized batch of examples, let's watch this video from the DeepLearning.AI course by Andrew Ng.

Watch the entire video. Focus on: The general equations for computing Z and A for any layer l (0:00 - 3:10). How the equations are adapted for a vectorized implementation across multiple training examples (3:10 - 5:25). The note that using a for loop to iterate through the layers is standard practice, even in vectorized implementations (5:25 - 6:30).

3. Building an MLP from Scratch with NumPy

Now for the most important part: translating theory into code. We'll build a simple two-layer neural network from scratch using only Python and NumPy to classify handwritten digits from the famous MNIST dataset. This is a classic "Hello, World!" for deep learning and will solidify your understanding of the mechanics.

The video below will be our guide. It walks through the entire process, from the theoretical setup to the final code.

Building a neural network FROM SCRATCH (no Tensorflow/Pytorch, just numpy & math)

This fantastic video by Samson Zhang will guide us through building a neural network from scratch. First, let's review the architecture and the specific activation functions he uses.

Watch from 02:14 to 07:56. Pay attention to: The Architecture: An input layer (784 nodes for 28x28 pixels), a hidden layer (10 nodes), and an output layer (10 nodes, one for each digit from 0-9). Forward Propagation Math: The equations for Z1, A1, Z2, and A2. Activation Functions: ReLU (Rectified Linear Unit) for the hidden layer. It's simple and effective: ReLU(z) = max(0, z). Softmax for the output layer. This is perfect for multi-class classification, as it converts the raw outputs (logits) into a probability distribution, where all outputs are between 0 and 1 and sum to 1.

Now that we have the plan, let's implement it. We will follow the video's coding part step-by-step.

Step 3.1: Data Loading and Preparation

First, we load the MNIST dataset and prepare it for our network. This involves splitting it into training and development sets and shaping the data so that each column represents one image.

Building a neural network FROM SCRATCH (no Tensorflow/Pytorch, just numpy & math)

Let's start by getting our data ready. The video uses pandas to load the CSV and numpy for all numerical operations.

Watch from 11:45 to 15:16. Notice the data transposition (.T) step. This is a common convention to structure the input matrix X with features as rows and examples as columns (shape n_features x n_examples).

Step 3.2: Initializing Parameters

Before we can do any computation, we need to initialize the weights () and biases (). We'll initialize them with small random values. This random start is crucial for the learning process, which we will see in the next lesson.

# Based on the video's implementation
import numpy as np

def init_params():
    # Layer 1: 10 neurons, 784 inputs
    W1 = np.random.rand(10, 784) - 0.5
    b1 = np.random.rand(10, 1) - 0.5
    # Layer 2: 10 neurons, 10 inputs
    W2 = np.random.rand(10, 10) - 0.5
    b2 = np.random.rand(10, 1) - 0.5
    return W1, b1, W2, b2

The init_params function creates the weight matrices and bias vectors for our two-layer network. The dimensions are determined by the number of neurons in the current and previous layers. For example, W1 connects 784 input nodes to 10 hidden neurons, hence its shape is (10, 784).

Step 3.3: Implementing Activation Functions

Next, we need to code our activation functions, ReLU and Softmax. Using NumPy, we can apply these functions efficiently across entire vectors or matrices.

def ReLU(Z):
    return np.maximum(Z, 0)

def softmax(Z):
    # The `axis=0` ensures we sum over the 10 outputs for each example in the batch.
    exp_Z = np.exp(Z)
    return exp_Z / np.sum(exp_Z, axis=0)

These functions implement the ReLU and Softmax activations. np.maximum provides an elegant element-wise implementation of ReLU. For Softmax, we exponentiate the inputs and then normalize them column-wise to get probabilities for each training example.

Step 3.4: Implementing Forward Propagation

Finally, we combine everything into our forward_prop function. This function will take the network parameters and an input batch and execute the sequence of calculations we defined earlier.

def forward_prop(W1, b1, W2, b2, X):
    # Layer 1
    Z1 = W1.dot(X) + b1
    A1 = ReLU(Z1)
    
    # Layer 2
    Z2 = W2.dot(A1) + b2
    A2 = softmax(Z2)
    
    return Z1, A1, Z2, A2

This function directly translates the forward propagation algorithm into code. It calculates the pre-activations (Z1, Z2) and activations (A1, A2) for both layers sequentially. Note the use of .dot() for matrix multiplication.

For a guided walkthrough of coding these functions, you can refer back to the video.

Test your understanding!

Suppose you are building an MLP with the following architecture:

  • Input layer: 256 features
  • Hidden layer 1: 128 neurons
  • Hidden layer 2: 64 neurons
  • Output layer: 10 neurons (for 10-class classification)

What are the shapes (dimensions) of the weight matrices , , and the bias vectors , , ? Assume the convention where .

Show answer
  • Layer 1 (Input -> Hidden 1):
    • must transform an input of size 256 to an output of size 128. Shape: (128, 256)
    • is a bias for each of the 128 neurons. Shape: (128, 1)
  • Layer 2 (Hidden 1 -> Hidden 2):
    • must transform an input of size 128 to an output of size 64. Shape: (64, 128)
    • is a bias for each of the 64 neurons. Shape: (64, 1)
  • Layer 3 (Hidden 2 -> Output):
    • must transform an input of size 64 to an output of size 10. Shape: (10, 64)
    • is a bias for each of the 10 neurons. Shape: (10, 1)

Conclusion

Congratulations on building your first neural network! Today, we've laid a critical foundation for the rest of our journey into deep learning.

Key Takeaways:

  • A Multi-Layer Perceptron (MLP) is a neural network made of an input layer, one or more hidden layers, and an output layer.
  • Each neuron computes a weighted sum of its inputs plus a bias, then passes the result through a non-linear activation function.
  • Forward propagation is the process of passing data through the network, layer by layer, to generate a prediction.
  • This process can be implemented efficiently using matrix operations, which is standard practice in all deep learning frameworks.
  • Activation functions like ReLU are used in hidden layers to enable learning of complex patterns, while functions like Softmax are used in the output layer to produce a probability distribution for classification.

Preview of the next lesson:
Our network can now make predictions, but they are completely random because the weights were initialized randomly. The next crucial question is: how does the network learn from its mistakes? In our next lesson, we will derive and implement the famous backpropagation algorithm, which is the mechanism that allows the network to calculate its errors and adjust its weights and biases to improve its predictions.

Can't find a good explanation? Sign up and we'll make it for you

Sign up