Hello! Welcome to the final lesson in our module on Mathematical and Statistical Foundations.
Introduction
In our last lesson, we explored the Law of Large Numbers and the Central Limit Theorem. These powerful theorems give us confidence that statistics calculated from a data sample (like the sample mean) are reliable estimators of the true, underlying population parameters. The LoLN assures us that our estimate gets better with more data, and the CLT tells us about the distribution of the estimation error.
This raises a crucial question: How do we find the "best" parameters for a chosen probability distribution to explain our data? This lesson directly addresses this by covering the learning outcome: Perform maximum likelihood estimation (MLE) and maximum a posteriori (MAP) estimation.
We will explore two fundamental principles for parameter estimation:
- Maximum Likelihood Estimation (MLE): This method asks, "What parameter values make the observed data most probable?" It is a cornerstone of frequentist statistics and implicitly underlies many machine learning loss functions.
- Maximum a Posteriori (MAP) Estimation: This is the Bayesian counterpart. It asks, "Given the data and my prior beliefs about the parameters, what are the most probable parameter values?" This approach allows us to incorporate prior knowledge, which can be incredibly powerful, especially with limited data.
Understanding MLE and MAP is not just a statistical exercise; it provides the theoretical justification for the objective functions we aim to optimize when training machine learning models.
1. The Core Idea: What Parameter Fits Best?
Imagine you have a set of data points. You hypothesize that they were generated from a Gaussian (normal) distribution, but you don't know the distribution's mean or variance . Which Gaussian distribution is the most "likely" one to have produced your data?

MLE provides a formal way to answer this. It directs us to choose the parameters (, ) that maximize the probability (or more accurately, the probability density) of observing our specific data points.
Let's watch a video that builds a fantastic visual intuition for this process.
What are Maximum Likelihood (ML) and Maximum a posteriori (MAP)? ("Best explanation on YouTube")
This video, 'What are Maximum Likelihood (ML) and Maximum a posteriori (MAP)?' from the channel 'Iain Explains Signals, Systems, and Digital Comms', provides an excellent intuitive walkthrough of MLE before introducing MAP.
Watch the segment from 01:51 to 10:26. The video uses a simple model where a measurement y is a scaled version of an unknown parameter x plus some Gaussian noise (y = ax + noise). Focus on how it visually demonstrates the search for the parameter x that maximizes the likelihood of having measured the specific value y_bar.
The key idea is that for each possible value of the parameter, we get a different probability distribution for the data. We then "reverse" the question: for the data we actually did observe, which parameter value makes that observation have the highest probability?
2. Maximum Likelihood Estimation (MLE)
Now, let's formalize the intuition from the video.
The Likelihood Function
Let's say we have a dataset of data points. We assume these points are independent and identically distributed (i.i.d.) and drawn from a probability distribution parameterized by .
Because the data points are independent, the joint probability of observing the entire dataset is the product of the individual probabilities:
This expression, when viewed as a function of the parameter for a fixed dataset , is called the Likelihood Function, .
The MLE principle states that the best estimate for is the one that maximizes this likelihood function:
The Log-Likelihood Trick
Working with products is mathematically cumbersome. Since the logarithm is a monotonically increasing function, maximizing a function is equivalent to maximizing its logarithm. This gives rise to the log-likelihood function, , which is much easier to work with.
This simple trick turns a product into a sum, which is far easier to differentiate—a key step in optimization. Your computer science background will appreciate that this also avoids numerical underflow that can occur when multiplying many small probabilities together.
The following video provides a more formal introduction to the likelihood function and a clear explanation of why we use the log-likelihood.
Maximum Likelihood Estimation (MLE) with Examples
Let's watch a segment from Steve Brunton's 'Maximum Likelihood Estimation (MLE) with Examples'. He provides a concise and mathematically grounded explanation of the likelihood and log-likelihood functions.
Watch from 01:01 to 12:27. This part defines the joint probability function, introduces the likelihood function, and crucially explains the motivation and benefit of using the log-likelihood for optimization.
Deriving an MLE: The Coin Flip Example
The classic example for MLE is estimating the bias of a coin. Let's say we flip a coin times and observe heads (outcome 1) and tails (outcome 0). We want to estimate . This is a Bernoulli distribution.
The following reading from Tom Mitchell's renowned machine learning book provides a rigorous yet clear derivation.
CHAPTER 2 Estimating Probabilities
This chapter excerpt from Tom Mitchell's 'Machine Learning' book formally derives the MLE for the coin-flipping problem. It's a foundational example that solidifies the concepts we've discussed.
Read Section 2.1, 'Maximum Likelihood Estimation (MLE)' (pages 8-10). The section starts with the principle, defines the likelihood for the Bernoulli distribution, and uses calculus (taking the derivative of the log-likelihood and setting it to zero) to derive the final estimate. Pay close attention to how the steps flow from the principle to the final, intuitive result.
As the text derives, the process is:
- Likelihood:
- Log-Likelihood:
- Differentiate:
- Set to zero and solve:
The result is exactly what our intuition would suggest: the best estimate for the probability of heads is the proportion of heads observed in the data. MLE gives us a formal justification for this intuitive answer.
3. Maximum a Posteriori (MAP) Estimation
MLE works well with lots of data. But what if our data is scarce?
Imagine you flip a new coin 3 times and get 3 heads.
The MLE estimate is .
This suggests the coin will always land on heads, which seems like an overconfident conclusion based on such little evidence. We have a prior belief that coins are usually fair, or at least close to fair.
MAP estimation allows us to incorporate such prior beliefs.
From Likelihood to Posterior
MAP shifts the goal from maximizing the likelihood to maximizing the posterior probability . We use Bayes' Theorem to define the posterior:
Here:
- is the posterior: what we believe about after seeing the data.
- is the likelihood: same as in MLE.
- is the prior: our belief about before seeing any data.
- is the evidence: the probability of the data.
The MAP estimate is the that maximizes this posterior:
Since is constant with respect to , it doesn't affect the location of the maximum. We can therefore ignore it for optimization:
In log space, this becomes:
This final expression is insightful: MAP estimate = MLE term + Log-Prior term. The prior acts as a "penalty" or a "regularizer" that pulls the estimate away from the pure MLE solution towards our prior belief.
Let's return to the video for a visual explanation of how the prior influences the estimate.
What are Maximum Likelihood (ML) and Maximum a posteriori (MAP)? ("Best explanation on YouTube")
We'll watch the second part of the 'Iain Explains' video, which introduces the MAP concept.
Watch from 10:26 to 18:06. This segment explains the need for a prior, defines the MAP estimate using Bayes' rule, and visually contrasts it with the MLE estimate, showing how the prior 'weights' the likelihood.
Deriving a MAP Estimate: Coin Flip with a Beta Prior
To perform MAP estimation, we must first define a prior distribution . For the coin flip example, is a probability between 0 and 1. A flexible distribution for this is the Beta distribution.
The Beta distribution, , is parameterized by two hyperparameters, and , which can be interpreted as representing prior knowledge equivalent to having seen "imaginary" heads and "imaginary" tails.
The Beta distribution is a conjugate prior to the Bernoulli likelihood. This is a special relationship: if the prior is Beta and the likelihood is Bernoulli, the resulting posterior is also a Beta distribution. This makes the math very convenient.
The same reading from Tom Mitchell's book now walks us through the MAP derivation.
CHAPTER 2 Estimating Probabilities
We will now see the formal derivation for the MAP estimate, which parallels the MLE derivation but includes the prior.
Read Section 2.2, 'Maximum a Posteriori Probability Estimation (MAP)' (pages 10-12). This section introduces the MAP principle, uses the Beta distribution as a prior, and derives the final MAP estimate. Notice how the prior parameters eta_1 and eta_0 are simply added to the observed data counts.
The derived MAP estimate is:
Going back to our 3-flips-3-heads example:
- .
- .
- Let's assume a prior belief that the coin is fair. We can represent this with (representing 1 imaginary head and 1 imaginary tail).
- .
The MAP estimate is much more reasonable than the MLE of 1.0. As we collect more data, the observed counts and will dwarf the small prior counts, and the MAP estimate will converge to the MLE estimate.
Test your understanding!
You are A/B testing a new "Add to Cart" button. You show it to 20 users. 15 of them click it, and 5 do not.
- What is the Maximum Likelihood Estimate for the click-through rate (CTR) of this button?
- Based on historical data for similar buttons, your team believes the CTR should be around 60%. You model this prior belief using a Beta distribution with hyperparameters and . (These correspond to prior "clicks" and prior "non-clicks", reflecting a 60% rate). What is the MAP estimate for the CTR?
- Which estimate is higher? Why?
Show answer
-
MLE Estimate:
The MLE for a Bernoulli process is simply the ratio of successes to total trials.- Observed clicks (): 15
- Observed non-clicks (): 5
- .
The MLE is a CTR of 75%.
-
MAP Estimate:
We use the MAP formula, incorporating the prior counts from the Beta distribution's hyperparameters.- Prior "clicks" ():
- Prior "non-clicks" ():
- .
The MAP estimate is a CTR of 70%.
-
Comparison:
The MLE estimate (75%) is higher than the MAP estimate (70%). This is because the observed data (75% CTR) had a higher rate than our prior belief (60% CTR). The MAP estimate is a compromise, "pulled" from the MLE result toward the prior.
4. Connection to Machine Learning
The principles of MLE and MAP are central to training machine learning models.
-
Loss Functions from MLE: Many common loss functions are simply the negative log-likelihood of the model.
- Mean Squared Error (MSE) for regression is the negative log-likelihood of the data assuming the target variable is the model's prediction plus some Gaussian noise.
- Cross-Entropy Loss for classification is the negative log-likelihood assuming the data follows a Bernoulli or Categorical distribution.
- Therefore, minimizing these loss functions is equivalent to performing Maximum Likelihood Estimation of the model's parameters.
-
Regularization from MAP: When you add a regularization term to your loss function, you are often implicitly performing MAP estimation.
- L2 Regularization (e.g.,
loss + λ||w||²) is equivalent to MAP estimation where you've placed a Gaussian prior with mean 0 on the model's weightsw. This penalizes large weights, assuming a priori that most weights should be small. - L1 Regularization (e.g.,
loss + λ||w||₁) is equivalent to MAP estimation with a Laplace prior on the weights. This prior is sharply peaked at zero, encouraging many weights to become exactly zero (sparsity).
- L2 Regularization (e.g.,
This provides a powerful probabilistic interpretation for techniques you'll frequently use in practice.
Conclusion
This lesson completes our foundational module by providing the principles for fitting models to data.
Key Takeaways:
- Maximum Likelihood Estimation (MLE) finds the parameter that maximizes the likelihood of the observed data, . It's a powerful principle but can overfit on small datasets.
- Maximum a Posteriori (MAP) Estimation finds the that maximizes the posterior probability . It incorporates a prior belief and is equivalent to adding a regularization term to the likelihood.
- The log-likelihood is a practical tool that turns products into sums, simplifying the math and improving numerical stability.
- Minimizing standard loss functions (like MSE or Cross-Entropy) is equivalent to MLE. Adding regularization (like L1 or L2) is equivalent to MAP.
Preview of the next lesson:
We have now established the statistical principles for what we want to optimize (the likelihood or the posterior). In the next module, "Core Machine Learning Concepts," we will dive into how we perform this optimization for complex models where we can't simply solve for the parameters analytically. We will begin with Gradient Descent, the fundamental algorithm used to train almost all modern neural networks by iteratively minimizing a loss function.