Hello! Let's dive into the next lesson in our journey through the mathematical foundations of AI.
In the previous lesson, we established a vocabulary for describing uncertainty using probability distributions like the Bernoulli, Categorical, and Gaussian. We saw that these distributions are defined by parameters, such as the mean () and variance () for a Gaussian. But what exactly are these parameters, and how do we compute them for any given random variable?
This lesson addresses the learning outcome: Compute expectations and variances of random variables. We will formalize the concepts of "center" and "spread" for probability distributions. The expected value will give us the theoretical average outcome, while the variance will quantify how spread out the outcomes are around this average. These two measures are fundamental for summarizing and understanding the behavior of random variables in virtually every area of machine learning.
1. Expected Value: The Center of Mass
The expected value, or expectation, of a random variable is its theoretical mean. It's the value you would expect to get on average if you could repeat an experiment an infinite number of times. It's denoted as or .
Think of it as a weighted average of all possible outcomes, where each outcome is weighted by its probability.

Calculating Expected Value for Discrete Variables
For a discrete random variable that can take values with probabilities , the expected value is the sum of each value multiplied by its probability:
Let's watch a video that walks through this definition and a clear example.
The Expected Value and Variance of Discrete Random Variables
This video from the jbstatistics channel, 'The Expected Value and Variance of Discrete Random Variables', provides an excellent, clear explanation of expected value and how to calculate it for a discrete random variable.
Watch the segments from 00:43 to 02:24 and 04:27 to 06:08. The first part defines expected value, and the second part applies the formula to a concrete example of a biased coin toss.
From Theoretical Expectation to Sample Mean
It's important to distinguish between the theoretical expected value of a random variable's distribution and the sample mean of a dataset.
- Expected Value : A property of the probability distribution itself. It's a constant.
- Sample Mean : An estimate of the expected value calculated from a finite sample of data. It's a variable that changes depending on the sample.
Given your background in software engineering, you've likely used functions to calculate the mean of a list of numbers. In Python with NumPy, this is straightforward.
A Gentle Introduction to Expected Value, Variance, and Covariance with NumPy
This article, 'A Gentle Introduction to Expected Value, Variance, and Covariance' by Machine Learning Mastery, provides a clear definition and practical NumPy examples.
Read the section 'Expected Value'. Pay attention to the distinction between the probability-weighted sum (theoretical expectation) and the simple average (sample mean). Also, review the Python code showing how to use numpy.mean().
The numpy.mean() function calculates the sample mean. The Law of Large Numbers (which we will cover in the next lesson) formally states that as your sample size grows, the sample mean will converge to the theoretical expected value.
Calculating Expected Value for Continuous Variables
For a continuous random variable, the concept is the same, but the summation is replaced by an integral over the probability density function (PDF), :
This integral represents the "center of mass" of the area under the PDF curve.
Deriving the Mean and Variance of a Continuous Probability Distribution
Let's watch another video from jbstatistics, this time demonstrating the calculation for a continuous random variable.
Watch from 00:43 to 03:01. This section shows how to set up and solve the integral to find the expected value for a given PDF. Your experience with calculus from your engineering studies should make this familiar.
2. Variance: The Spread of Outcomes
While expected value tells us about the center of a distribution, it doesn't tell us anything about how spread out the values are. That's the job of variance.
Variance, denoted as or , measures the average squared difference of a random variable from its mean. A small variance means the outcomes are tightly clustered around the expected value, while a large variance means they are widely scattered.

Calculating Variance
The definition of variance is the expectation of the squared deviation from the mean :
For discrete variables, this becomes:
For continuous variables, this becomes:
While this formula is definitionally pure, a more computationally convenient formula exists:
This formula states that the variance is the "mean of the square minus the square of the mean." To use it, you first calculate and , then plug them in.
Let's see how both formulas are used in practice.
The Expected Value and Variance of Discrete Random Variables
We'll return to the jbstatistics videos, which do an excellent job of demonstrating the calculation of variance using both the definitional and computational formulas.
First, watch the explanation for the discrete case from 02:50 to 04:12 and 07:26 to 11:10. This covers the two formulas and then applies both to the coin toss example, showing they yield the same result.
Deriving the Mean and Variance of a Continuous Probability Distribution
Now, let's see the same concept applied to the continuous case.
Watch from 03:01 to 07:17. This section walks through calculating E[X²] via integration and then uses the computational formula Var(X) = E[X²] - (E[X])² to find the variance.
Standard Deviation
The standard deviation, denoted , is simply the square root of the variance:
The primary advantage of using standard deviation is that it has the same units as the random variable , making it more interpretable as a typical amount of deviation from the mean.
As with the mean, NumPy provides a simple way to calculate the variance and standard deviation of a sample of data.
A Gentle Introduction to Expected Value, Variance, and Covariance with NumPy
Let's revisit the Machine Learning Mastery article to see the practical implementation.
Read the section 'Variance'. Note the ddof (Delta Degrees of Freedom) parameter in numpy.var(). Setting ddof=1 calculates the sample variance, which is an unbiased estimator for the population variance. This is a subtle but important point in statistics.
Test your understanding!
You are playing a game with a single six-sided die. However, the die is loaded. The probability of rolling each number is as follows:
| Outcome (X) | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| P(X=x) | 0.1 | 0.1 | 0.1 | 0.1 | 0.1 | 0.5 |
Calculate the expected value and the variance of a single roll.
Show answer
1. Calculate the Expected Value :
We use the formula .
The expected value of a roll is 4.5.
2. Calculate the Variance :
We'll use the computational formula . First, we need to find .
Now, we plug this into the variance formula:
The variance of a roll is 3.25. The standard deviation would be .
3. Properties of Expectation and Variance
Understanding how expectation and variance behave under transformations is crucial in machine learning. Let be a random variable and be constants.
ECE 595: Machine Learning I Tutorial 02: Probability Review
This short review from Purdue University provides a concise summary of the key properties of expectation and variance.
Review slides 12 ('Properties of Expectation') and 13 ('Properties of Variance'). Focus on understanding the rules for scaling and shifting a random variable.
Here's a summary of the most important properties:
Properties of Expectation:
- Linearity: . This is highly intuitive. If you scale a variable and shift it, its average scales and shifts the same way.
- Sum of Variables: . The expectation of a sum is the sum of the expectations. This holds even if and are dependent.
Properties of Variance:
- DC Shift: . Adding a constant shifts the entire distribution but does not change its spread. The variance is unchanged.
- Scale: . Scaling a variable by scales its variance by . The square is important; since variance is in squared units, the scaling factor is also squared.
- Combined: .
These properties are used extensively in proofs and derivations for ML algorithms, such as analyzing the effects of batch normalization or weight initialization.
Conclusion
In this lesson, we moved from describing distributions to summarizing them numerically. You learned how to compute the two most important summary statistics for any random variable.
Key Takeaways:
- Expected Value () is the probability-weighted average of a random variable's outcomes, representing its center or theoretical mean.
- Variance () is the expected squared deviation from the mean, quantifying the spread or dispersion of the outcomes.
- The calculation method depends on the variable type: summation for discrete variables and integration for continuous variables.
- A computationally efficient formula for variance is .
- Expectation is a linear operator (), while variance is not ().
Preview of the next lesson:
We have now established a clear difference between the theoretical mean () of a distribution and the sample mean () calculated from data. But how are they related? How can we be sure that our sample mean is a good estimate of the true, underlying theoretical mean? The next lesson will answer this by exploring two cornerstone theorems of statistics: the Law of Large Numbers and the Central Limit Theorem. These theorems provide the formal justification for why we can use statistics from a sample to make inferences about an entire population or distribution.