Hello! Let's get started with today's lesson.
Introduction
In our previous lessons, we've focused on building individual predictive models like Decision Trees and K-Nearest Neighbors. Each of these algorithms learns from data in its own unique way to make predictions. However, a powerful concept in machine learning is that a committee of models can often make better decisions than any single expert. This is the core idea behind ensemble learning.
Today, we will explore the foundational principles of this approach, focusing on two of the most influential techniques: Bagging and Boosting. This lesson directly addresses the learning outcome: Explain the principles of ensemble learning, including bagging and boosting.
We will cover:
- The fundamental concept of ensemble learning.
- Bagging (Bootstrap Aggregating): How training models in parallel on different data samples can reduce variance and improve stability.
- Boosting: How training models sequentially, with each one correcting its predecessor's mistakes, can reduce bias and build a highly accurate predictor.
- A direct comparison to clarify the key differences, strengths, and use cases of each technique.
Understanding these principles is crucial as they form the basis for some of the most successful algorithms in machine learning, such as Random Forests and Gradient Boosting Machines, which we will implement in upcoming lessons.
1. The "Wisdom of the Crowd": What is Ensemble Learning?
At its heart, ensemble learning is about combining the predictions of several individual models (often called "base learners" or "weak learners") to produce a single, superior "strong learner". The intuition is that by aggregating multiple perspectives, we can average out the individual errors, leading to a final prediction that is more accurate and robust.
This short video provides a great high-level introduction to the concept.
Ensemble (Boosting, Bagging, and Stacking) in Machine Learning: Easy Explanation for Data Scientists
Watch this introductory segment from 'Ensemble (Boosting, Bagging, and Stacking) in Machine learning' by Emma Ding to grasp the main idea.
Watch from the beginning to 01:27. Focus on how a group of weak learners can form a strong learner and why this makes the final model more robust and less prone to overfitting.
As the video explains, if one model makes an error, other models in the ensemble can correct for it. This collective decision-making process is what gives ensemble methods their power. Now, let's explore the two primary ways to build such an ensemble.
2. Bagging: Reducing Variance through Parallelism
Bagging, which stands for Bootstrap Aggregating, is an ensemble technique that aims to reduce the variance of a predictive model. High variance is a sign of overfitting, where a model learns the training data too well, including its noise, and fails to generalize to new data.
Bagging tackles this with a two-step process:
- Bootstrap: It creates numerous different training datasets from the original one. It does this by sampling with replacement. Imagine you have a bag of 100 data points. You'd create a new dataset by randomly picking a point, noting it down, and putting it back in the bag. You repeat this 100 times. The resulting dataset will have the same size as the original, but some points will be duplicated, and others will be missing entirely. This process is repeated many times to create multiple, diverse training sets.
- Aggregate: A base learner (like a decision tree) is trained independently on each of these bootstrapped datasets. Since the models are independent, this can be done in parallel. To get a final prediction, their results are aggregated:
- For classification, a majority vote is taken.
- For regression, the predictions are averaged.
By averaging the outputs of many models trained on slightly different data, the "noise" and instability (variance) of any single model are smoothed out.
This video from StatQuest vividly illustrates the bootstrapping and aggregation process in the context of Random Forests, which are a direct application of bagging.
StatQuest: Random Forests Part 1 - Building, Using and Evaluating
Watch these segments from 'StatQuest: Random Forests Part 1' to see a practical demonstration of bootstrapping and aggregation (bagging).
Focus on the mechanics of the process: Creating Bootstrap Datasets (01:21 - 02:24): Observe how new datasets are created by sampling with replacement. Building Multiple Trees (04:00 - 04:34): Understand that a separate model is built on each bootstrap sample. Making a Prediction (04:34 - 05:53): See how the final prediction is made through voting. The video explicitly labels this entire process as bagging at the end.
The Impact of Bagging on Bias and Variance
Bagging is most effective when used with models that have low bias but high variance. A classic example is a fully grown decision tree, which can perfectly fit the training data (low bias) but is very sensitive to small changes in it (high variance).
Let's consider the math. The variance of the average of independent random variables is times the variance of a single variable. While the predictions of our bootstrapped models aren't perfectly independent (since their training sets overlap), they are sufficiently decorrelated for the averaging process to significantly reduce the overall variance. The bias, however, remains largely unchanged because the expectation of the average prediction is the same as the expectation of a single prediction.
Bagging, Boosting, and Stacking in Machine Learning
For a concise summary of the bagging process, its goals, and its pros and cons, please read this section from Baeldung's article on ensemble models.
Read 'Section 2. Bagging', including the subsections '2.1. Algorithms That Use Bagging' and '2.2. Pros and Cons of Bagging'. Focus on the core steps and the main goal of reducing variance.
3. Boosting: Reducing Bias through Sequential Learning
Boosting takes a completely different approach. Instead of building models in parallel, it builds them sequentially, where each new model attempts to fix the errors made by the previous ones. The goal of boosting is primarily to reduce a model's bias. High bias indicates underfitting, where a model is too simple to capture the underlying patterns in the data.
The process generally works as follows:
- A simple base model (a "weak learner," e.g., a decision tree with only one split, called a "stump") is trained on the data.
- The predictions are evaluated. The data points that were misclassified are given higher weights.
- A second weak learner is trained, but this time it pays more attention to the highly-weighted (previously incorrect) data points.
- This process repeats for a specified number of iterations. Each successive learner focuses on the "hardest" examples that the ensemble is still struggling with.
- To make a final prediction, the predictions from all learners are combined in a weighted sum. Learners that performed better (had lower error on their respective weighted datasets) are given more say in the final decision.

The Impact of Boosting on Bias and Variance
Boosting combines many high-bias, low-variance models (weak learners) to create a single, highly complex strong learner with low bias. By iteratively focusing on mistakes, the ensemble gradually learns the complex decision boundaries required to correctly classify the data.
Let's watch a video that visualizes this sequential process.
Ensemble (Boosting, Bagging, and Stacking) in Machine Learning: Easy Explanation for Data Scientists
This segment from Emma Ding's video clearly explains the sequential nature of boosting and its effect on bias and variance.
Watch the section on boosting from 03:40 to 05:44. Pay close attention to how the weights of misclassified examples are adjusted and how this process helps reduce the bias of the overall model.
Famous boosting algorithms like AdaBoost (Adaptive Boosting) and Gradient Boosting follow this core principle. While bagging is about getting a consensus from a diverse group of independent experts, boosting is like a team of specialists learning from each other's mistakes in a focused, iterative fashion.
Bagging, Boosting, and Stacking in Machine Learning
For a textual explanation of the boosting process and its characteristics, please read the corresponding section from the Baeldung article.
Read 'Section 3. Boosting', including its subsections. This will reinforce the sequential training concept and the goal of reducing bias.
4. Bagging vs. Boosting: A Summary
While both are powerful ensemble techniques, their philosophies and applications are distinct. The choice between them often depends on the nature of the problem and the base models you are working with.

Here is a summary of the key differences:
| Criteria | Bagging | Boosting |
|---|---|---|
| Approach | Trains base models in parallel. | Trains base models sequentially. |
| Model Dependency | Models are independent of each other. | Each model depends on the previous ones. |
| Primary Goal | To reduce variance (combat overfitting). | To reduce bias (combat underfitting). |
| Base Learners | Works best with complex, high-variance models (e.g., deep decision trees). | Works best with simple, high-bias "weak" learners (e.g., decision stumps). |
| Data Weighting | All training samples have equal weight (via bootstrap sampling). | Misclassified samples are given higher weights for the next model. |
| Prediction | Combines predictions via simple averaging or majority voting. | Combines predictions via a weighted average or vote. |
Test your understanding!
You are tasked with improving the performance of a machine learning system for a competition. You have two initial observations:
- Your first model is a single, very deep decision tree. It achieves 99.9% accuracy on the training set but only 75% on the test set.
- Your second attempt uses a collection of very simple linear models, each of which only achieves about 55% accuracy on the training set (where 50% is random chance).
Which ensemble technique, bagging or boosting, would be a better choice to improve each of these models, and why?
Show answer
-
The deep decision tree is suffering from high variance (overfitting). It has learned the training data perfectly but doesn't generalize. Bagging is the ideal choice here. By training many deep trees on different bootstrap samples and averaging their results, bagging will smooth out the decision boundary and drastically reduce the variance, leading to better generalization on the test set.
-
The simple linear models are suffering from high bias (underfitting). They are too simple to capture the complexity of the data. Boosting is the right tool for this job. It is specifically designed to take a collection of "weak learners" (models that are only slightly better than random) and sequentially combine them into a single, powerful "strong learner" with low bias.
Conclusion
In this lesson, we have demystified the core principles of ensemble learning by focusing on bagging and boosting. You now understand how these two powerful meta-algorithms leverage the "wisdom of the crowd" in fundamentally different ways to create models that are often far superior to their individual components.
Key Takeaways:
- Ensemble learning combines multiple models to improve predictive performance and robustness.
- Bagging is a parallel method that trains models on bootstrapped data subsets. Its main purpose is to reduce variance. It works best with complex, unstable models.
- Boosting is a sequential method where each model focuses on correcting the errors of its predecessor. Its main purpose is to reduce bias. It works by combining many simple, weak learners.
- The choice between bagging and boosting depends on whether the primary problem is high variance (overfitting) or high bias (underfitting).
Preview of the next lesson:
Now that you have a solid theoretical foundation in bagging, we are ready to get our hands dirty. In the next lesson, we will implement one of the most popular and effective machine learning algorithms in existence: the Random Forest. You will see how it combines the decision trees we've already built with the bagging principles you learned today.