Skip to main content
Create your own

Common Probability Distributions and Their Applications

Hello! Welcome to your next lesson on the mathematical foundations of AI.

In our last session, we explored Bayes' theorem and saw how it provides a formal framework for updating our beliefs in light of new evidence. We ended with a key question: how do we mathematically represent these beliefs and likelihoods, like or ? The answer lies in probability distributions.

This lesson will address the learning outcome: Describe common probability distributions (Gaussian, Bernoulli, Categorical) and their applications. These three distributions are fundamental building blocks in machine learning. You'll find them everywhere, from defining the outputs of classifiers to modeling noise in your data. Understanding them is crucial for building and interpreting AI models.

1. What is a Probability Distribution?

At its core, a probability distribution is a mathematical function that describes the likelihoods of different possible outcomes for a random variable.

For a deeper look into the formal definitions, please read the first two sections of the following article.

Probability Distributions in Machine Learning

This article from 'illumination' on Medium, titled 'Probability Distributions in Machine Learning', clearly defines what a probability distribution is and distinguishes between discrete and continuous variables.

Read the sections 'What is a Probability Distribution?' and 'Why They Matter in Machine Learning'. Focus on understanding the distinction between a Probability Mass Function (PMF) for discrete variables and a Probability Density Function (PDF) for continuous variables.

To summarize the key ideas:

  • Random Variable (X): A variable whose value is a numerical outcome of a random phenomenon.
  • Discrete Random Variable: Can only take on a countable number of distinct values (e.g., the result of a dice roll, the number of clicks on an ad).
  • Continuous Random Variable: Can take on any value within a given range (e.g., a person's height, the temperature of a room).
  • Probability Mass Function (PMF): For a discrete variable, the PMF, , gives the probability that the variable takes on a specific value . The sum of probabilities over all possible values is 1.
  • Probability Density Function (PDF): For a continuous variable, the PDF, , describes the relative likelihood of the variable taking on a value near . The probability of the variable falling within a specific range is the integral (area under the curve) of the PDF over that range. The total area under the curve is 1.

Now, let's explore the three key distributions for this lesson.

2. Discrete Distributions: Modeling Choices

Discrete distributions are used to model outcomes that fall into distinct categories.

The Bernoulli Distribution: A Single Yes/No Decision

The Bernoulli distribution is the simplest of all. It models a single experiment with only two possible outcomes, typically labeled as "success" (1) and "failure" (0).

To build your intuition, let's watch a short video segment.

10 Probability Distributions in 50 Min

This video, '10 Probability Distributions in 50 Min', provides a very intuitive, ground-up explanation of fundamental distributions. We'll start with its description of the Bernoulli distribution.

Watch from 04:58 to 07:06. This part explains the Bernoulli distribution as the 'atom of decision', defined by a single parameter P, the probability of success.

As the video explained, the Bernoulli distribution is defined by a single parameter, , which is the probability of success. Its PMF can be written elegantly as:

If (success), the formula becomes . If (failure), it becomes .

Application in Machine Learning: Binary Classification
The Bernoulli distribution is the probabilistic foundation for binary classification. When a model like logistic regression outputs a probability (e.g., "0.8 chance of being spam"), it's providing the parameter for a Bernoulli trial. The actual outcome (spam or not spam) is the result of that trial.

This connection is crucial for understanding why certain loss functions are used.

Maximum Likelihood Estimation

The following article, 'Maximum Likelihood Estimation - how neural networks learn', bridges the gap between probability theory and neural network training. It shows how assuming a data distribution leads directly to the loss functions we use.

Read the sections 'Bernoulli distribution' (under 'The likelihood function') and 'Bernoulli distribution: binary classifiers' (under 'Training neural networks'). Notice how minimizing the negative log-likelihood of a Bernoulli distribution results in the binary cross-entropy loss function, a standard in classification tasks.

This is a powerful insight: training a binary classifier by minimizing cross-entropy is equivalent to finding the model parameters that make the observed labels most likely, assuming the outcomes are generated from a Bernoulli distribution.

The Categorical Distribution: A Single Choice Among Many

What if an experiment has more than two outcomes, like rolling a die? This is where the Categorical distribution comes in. It generalizes the Bernoulli distribution for a single trial with possible outcomes.

Basic Probability Distributions Explained: Bernoulli, Binomial, Categorical, Multinomial

This video, 'Basic Probability Distributions Explained', gives a concise explanation of the Categorical distribution as a generalization of the Bernoulli for more than two outcomes.

Watch from 04:22 to 05:46. Pay attention to how it's defined by a vector of probabilities, one for each of the K categories.

The Categorical distribution is parameterized by a vector , where is the probability of the -th outcome, and .

Application in Machine Learning: Multi-class Classification
This is the backbone of multi-class classification. The softmax function at the end of a neural network for a task like image classification (e.g., identifying digits 0-9) outputs a probability vector. This vector parameterizes a Categorical distribution, giving the probability for each class.

Test your understanding!

You are building a machine learning model for two different tasks:

  1. Predicting whether a user will click an "Add to Cart" button (a single yes/no event).
  2. Predicting which genre of movie (e.g., 'Comedy', 'Drama', 'Sci-Fi', 'Horror') a user will rate highly.

Which probability distribution (Bernoulli or Categorical) would you use to model the outcome of each task?

Show answer
  1. Click Prediction: This is a binary outcome (click or no-click), so it's modeled by the Bernoulli distribution.
  2. Genre Prediction: This is a choice among multiple distinct categories, so it's modeled by the Categorical distribution.

3. Continuous Distributions: Modeling Measurements

Now let's turn to variables that can take any value within a range, like measurements of height, weight, or temperature.

The Normal (Gaussian) Distribution: The Bell Curve

The Normal, or Gaussian, distribution is arguably the most important distribution in all of statistics and machine learning. Its iconic "bell curve" shape appears everywhere.

Let's return to the "10 Distributions" video to see why it's so common and how to interpret it.

10 Probability Distributions in 50 Min

The video provides a great intuitive explanation for why the Normal distribution is so ubiquitous, connecting it to the Central Limit Theorem, and clarifies the roles of its two parameters: the mean and standard deviation.

Watch from 13:14 to 20:23. Focus on the concept of the Central Limit Theorem (why adding up random events leads to a bell curve) and how the mean (\mu) and standard deviation (\sigma) control the center and spread of the curve.

As the video explains, the Normal distribution is defined by two parameters:

  • Mean (): The center of the distribution; its peak.
  • Standard Deviation (): The measure of the spread or width of the distribution. A small means the data is tightly clustered around the mean, while a large means it's spread out. The variance is .

The PDF is given by the formidable-looking formula:

You don't need to memorize this. The key is to understand that it's a machine for drawing a bell curve centered at with a width determined by .

Application in Machine Learning: Regression and Initialization
The Normal distribution is central to regression tasks. We often assume that the errors (or "residuals") between our model's predictions and the actual continuous values are normally distributed.

Let's see how this assumption connects directly to the most common loss function for regression: Mean Squared Error (MSE).

Maximum Likelihood Estimation

We'll return to the 'Maximum Likelihood Estimation' article to see the connection between the Normal distribution and regression.

Read the sections 'Normal (Gaussian) distribution' (under 'The likelihood function') and 'Normal distribution: least squares regression' (under 'Training neural networks'). This will show you that minimizing the MSE is equivalent to finding the most likely parameters for a model whose errors are assumed to be Gaussian.

This is another profound connection. When you train a regression model using MSE, you are implicitly making a powerful assumption: that your model's predictions are off by some random noise that follows a Gaussian distribution.

Flowchart of Common Probability Distributions and Their Relationships
This chart visually summarizes the distributions we've discussed and others. You can see how Bernoulli is a simple case that leads to the Binomial (counting successes in multiple Bernoulli trials), and how the Normal (Gaussian) distribution is a central hub for continuous data.

Conclusion

Today we've demystified three of the most important probability distributions in machine learning. They are not just abstract mathematical concepts; they are the language we use to define our models' assumptions about the data.

Key Takeaways:

  • Probability Distributions describe the likelihood of outcomes for a random variable, using a PMF for discrete values and a PDF for continuous values.
  • The Bernoulli Distribution models a single trial with two outcomes (0/1). In ML, it's the foundation of binary classification, and maximizing its likelihood leads to the cross-entropy loss function.
  • The Categorical Distribution generalizes Bernoulli to a single trial with K outcomes. In ML, it's used for multi-class classification, with the softmax function providing its parameters.
  • The Normal (Gaussian) Distribution is a continuous distribution defined by its mean () and standard deviation (). It's ubiquitous due to the Central Limit Theorem. In ML regression, assuming Gaussian errors leads directly to the Mean Squared Error (MSE) loss function.

Preview of the next lesson:
Now that we can describe the shape and form of randomness with distributions, we need tools to summarize their key properties. In the next lesson, we will learn how to compute expectations and variances of random variables. The expectation is the formal term for the average value or mean (), and the variance () quantifies the spread, giving us precise ways to characterize these distributions.

Can't find a good explanation? Sign up and we'll make it for you

Sign up