Hello! Let's dive into the next lesson in our "Mathematical and Statistical Foundations for AI" module.
In our last lesson, we established how to find the gradient (), the vector of partial derivatives that points in the direction of steepest ascent for a function. This gave us the core principle of gradient descent: to minimize a loss function, we repeatedly take small steps in the direction of the negative gradient.
However, the loss functions for neural networks are not simple expressions like . They are deeply nested composite functions. The loss is a function of the model's predictions, which are a function of the final layer's activations, which are themselves functions of weights and the previous layer's activations... all the way back to the input data.
Today's learning outcome is to apply the chain rule for differentiating composite functions in the context of backpropagation. This is the mechanism that allows us to efficiently calculate the gradient of that final, complex loss function with respect to any weight, even one buried deep in the network's earliest layers. This process, known as backpropagation, is the engine that drives learning in virtually all modern neural networks.
1. The Core Idea: Computational Graphs and the Chain Rule
Trying to write out the full derivative of a neural network's loss function as one giant equation is a path to madness. As a software engineer, you're used to breaking down complex problems into modular, manageable components. We can do the same for differentiation using a concept that will feel very natural to you: the computational graph.
A computational graph represents a function as a directed graph where nodes are variables (scalars, vectors, matrices) and edges are operations or functions.
Let's start with a very simple function: . We can break this down and represent it as a graph:
- An
addnode takesxandyas input and produces an intermediate result,q = x + y. - A
multiplynode takesqandzas input and produces the final output,f = q * z.

The process of computing the gradient unfolds in two passes:
- Forward Pass: We feed in the inputs and compute the output, storing the intermediate values at each node. For , this gives and .
- Backward Pass: We start at the end and work backward, calculating the gradient at each step using the chain rule. The goal is to find , , and .
The chain rule tells us how to "chain" these derivatives together. For example, to find the effect of x on f, we must consider how x affects q, and then how q affects f:
CS231n Winter 2016: Lecture 4: Backpropagation, Neural Networks 1
Let's watch Andrej Karpathy's classic lecture from Stanford's CS231n course. He provides an exceptionally clear, step-by-step walkthrough of the forward and backward passes on this simple graph. His explanation of 'local' and 'global' gradients is fundamental.
Watch from 01:25 to 12:16. First, he motivates why we need computational graphs for complex models. Then, he meticulously walks through the f=(x+y)z example. Pay close attention to how he computes the local derivatives for each 'gate' (like '+' and '*') and then uses the chain rule to propagate the gradient backward.
To reinforce this, you can read the corresponding section in the course notes, which presents the same example in a textual format.
Neural Networks: Backpropagation
This section from the CS231n notes provides a written summary of the concept we just saw in the video.
Read the section titled 'Compound expressions, chain rule, backpropagation'. It provides the code and diagram for the (x+y)z example, reinforcing how the chain rule links the gradients together.
2. Developing Intuition: Local Gates and Propagating Sensitivity
The key insight from the previous example is that backpropagation is a local process. Each "gate" (or node) in the graph is a self-contained unit.
- During the forward pass, a gate takes inputs and computes its output. It can also compute its local gradients—the derivatives of its output with respect to its inputs, completely unaware of the rest of the network. For our
q = x + ygate, the local gradients are and . - During the backward pass, the gate receives an "upstream" gradient—the gradient of the final loss with respect to the gate's output ().
- The gate's job is to take this upstream gradient and multiply it by its local gradients to compute the "downstream" gradients (). It then passes these downstream gradients to the gates that fed into it.
This can be thought of as gates communicating with each other. The upstream gradient tells a gate "how much the final output wants it to increase or decrease," and the gate uses its local knowledge to pass this message along to its own inputs.
Neural Networks: Backpropagation
The CS231n notes offer a fantastic intuitive explanation of this 'local communication' view of backpropagation.
Read the section 'Intuitive understanding of backpropagation'. It anthropomorphizes the gates in a helpful way, explaining how they use gradient signals to communicate.
Another powerful way to visualize this is by thinking in terms of "sensitivity." The chain rule helps us measure how sensitive the final cost function is to a small change in a parameter deep inside the network.
Backpropagation calculus | Deep Learning Chapter 4
The 3Blue1Brown video on backpropagation provides a superb visual intuition for this idea of sensitivity. Instead of a generic computational graph, it uses a simple neural network from the start.
Watch from the beginning (00:21) to 06:43. Notice how he breaks down the derivative of the cost C with respect to a weight w into a chain of three ratios: ∂C/∂a, ∂a/∂z, and ∂z/∂w. This perfectly illustrates the chain rule in action within a network context.
3. Backpropagation in Practice: Code Structure and Modularity
Now let's connect these ideas to how backpropagation is actually implemented. Your experience as a software engineer will make these concepts particularly clear, as they mirror common software design patterns.
Staged Computation
When implementing backpropagation, you don't write one massive derivative function. You structure the forward pass into intermediate stages, just like we did with q = x + y. Then, in the backward pass, you compute the gradients for each stage in reverse order. This modularity is essential for debugging and for building complex models.
Modularity and "Gates"
What constitutes a "gate" is up to you. You could break a sigmoid activation function, , into many small gates (negation, exponentiation, addition, inversion). Or, you could treat the entire sigmoid function as a single gate. As long as you know its local derivative (), you can backpropagate through it in one step. This modularity is precisely how deep learning frameworks are built—they provide a library of pre-defined "layers" or "ops" that know their own forward and backward passes.
Caching Forward Pass Values
Notice that to compute the backward pass, you often need values computed during the forward pass. For example, in the f=qz gate, to compute , you need the value of q from the forward pass. Efficient implementations therefore cache these values during the forward pass to have them ready for the backward pass.
Gradient Accumulation
What happens if a variable is used in multiple parts of the graph? For example, in , the variable x "forks" and is used twice. The multivariable chain rule states that the gradients from all paths are simply added together. In implementation, this means when you backpropagate to x, you must use an += operation to accumulate the gradient, not = which would overwrite the previous value.
CS231n Winter 2016: Lecture 4: Backpropagation, Neural Networks 1
Let's return to Karpathy's lecture, where he discusses exactly these practical considerations. He shows a more complex sigmoid neuron example and then talks about how this translates into the API design of real deep learning frameworks.
Watch from 16:00 to 43:23. He first walks through backpropagation on a sigmoid neuron, showing the modularity. Then, he discusses the forward() and backward() API that every layer in a framework like PyTorch or Caffe (an earlier framework) implements. This directly connects the theory to the code you would write or use.
Test your understanding!
Consider a multiplication gate f(x,y) = x * y. During the forward pass, we find that x = 3 and y = -4. During the backward pass, the gate receives an upstream gradient . What are the downstream gradients and that this gate will pass backward?
Show answer
-
Find local gradients:
-
Apply the chain rule:
The downstream gradients are 8 for the x-path and -6 for the y-path.
4. A Glimpse into Reality: Vectorized Operations
So far, we've mostly used scalars. In practice, neural networks operate on vectors, matrices, and tensors. The math remains the same, but the derivatives become Jacobian matrices (which we will explore more formally in the next lesson).
For a function with a vector input and vector output , the "local derivative" is an Jacobian matrix where each entry is .
The chain rule becomes a matrix-vector product. However, a crucial point for efficiency is that we almost never explicitly compute and store these full Jacobian matrices. They are often huge and very sparse (mostly zeros).
Consider the ReLU activation function, , applied element-wise to a vector. The Jacobian matrix is diagonal because only affects . The diagonal entries are 1 if and 0 if . Instead of building this giant diagonal matrix and multiplying by it, we just implement the backward pass as an element-wise operation: pass the upstream gradient through where the input was positive, and block it (set to zero) where the input was negative.
CS231n Winter 2016: Lecture 4: Backpropagation, Neural Networks 1
Karpathy concludes his lecture by explaining this exact point about vectorized operations and why forming full Jacobians is impractical.
Watch from 43:23 to 47:40. This section explains how backpropagation extends to vectors and why the sparsity of Jacobians in common operations allows for very efficient implementations without forming the full matrix.
Conclusion
Today we've demystified backpropagation, revealing it to be a clever, recursive application of the chain rule on a computational graph. You now understand not just the mathematical formula, but also the computational paradigm that makes training deep networks possible.
Key Takeaways:
- Backpropagation is an algorithm for efficiently computing gradients of complex, composite functions by applying the chain rule backward through a computational graph.
- The process is local and modular: each operation ("gate") only needs to know its own local derivative and the upstream gradient it receives.
- Practical implementation involves patterns familiar from software engineering, such as staged computation, caching intermediate values, and accumulating results (
+=) when a variable is used in multiple places. - In real-world applications with vectors and matrices, we typically implement the effect of multiplying by the Jacobian matrix directly, rather than forming the large, sparse matrix itself.
Preview of the next lesson:
We've seen that the gradient (a vector of first-order derivatives) is essential for optimization. But what if we want to know about the curvature of the loss function? This requires second-order derivatives. In the next lesson, we will formally define Jacobians (matrices of first-order partials) and Hessians (matrices of second-order partials), exploring how they provide deeper insight into the optimization landscape and enable more advanced optimization algorithms.