Hello! Welcome to your next lesson.
Introduction
In our last lesson, we explored the foundational principles of ensemble learning, focusing on bagging and boosting. You learned that bagging (Bootstrap Aggregating) is a powerful technique for reducing a model's variance by training multiple independent models in parallel on different subsets of the data and then aggregating their predictions.
Today, we will put that theory into practice. This lesson is designed to help you implement a random forest classifier using bagging. A Random Forest is one of the most widely used and effective machine learning algorithms, and it's a direct and elegant application of the bagging principle with a clever twist.
As a brief recap, recall these key points from our last session:
- Bagging creates diverse training sets via bootstrapping (sampling with replacement).
- It trains a base model (like a decision tree) on each bootstrapped set.
- It combines their predictions through majority voting (for classification) or averaging (for regression).
Building on this, we will dive into the specifics of the Random Forest algorithm and implement it from scratch in Python. Given your software engineering background and familiarity with Python, you'll be well-equipped to construct the algorithm's architecture step-by-step.
1. What Makes a Forest "Random"?
A Random Forest is an ensemble of decision trees. We could simply apply bagging to a set of deep decision trees, and that would work quite well. However, this approach has a potential weakness: if the dataset has a few very strong, dominant features, most of the trees in the ensemble will likely select these same features for their initial splits. This makes the trees structurally similar and highly correlated. When predictions from correlated models are averaged, the variance reduction is less effective.
The Random Forest algorithm introduces a second layer of randomness to solve this very problem, further decorrelating the trees.

The two sources of randomness are:
- Bagging (Data Randomness): Each tree is trained on a different bootstrap sample of the original data. This is the core principle we learned in the previous lesson.
- Feature Subsampling (Feature Randomness): This is the key addition. When building each tree, at every split point, the algorithm does not consider all available features. Instead, it selects a random subset of features and only considers those for finding the best split.

This simple trick is incredibly effective. It prevents individual trees from becoming overly reliant on the same set of dominant features, forcing them to explore a wider variety of predictive patterns.
To solidify this concept, let's watch a short segment from a lecture that defines Random Forests.
Machine Learning Lecture 31 "Random Forests / Bagging" -Cornell CS4780 SP17
This video from Kilian Weinberger at Cornell University explains how Random Forests extend the idea of bagging by adding random feature selection.
Watch from 05:23 to 10:20. Pay attention to: How feature subsampling is introduced at each split in the decision tree. The reason for doing this: to make the individual trees even more different from each other (decorrelate them). The practical rule of thumb for choosing the number of features to sample (k or m): for a dataset with d features, it's often set to sqrt(d).
2. Implementing a Random Forest from Scratch
Now it's time to build our RandomForest classifier. The overall structure will be a class that manages a collection of DecisionTree objects. We will assume you have a working DecisionTree implementation from our previous lessons, as the Random Forest acts as a wrapper around it.
The following video provides a clear, step-by-step guide to implementing a Random Forest in Python. We will follow its structure to build our classifier.
How to implement Random Forest from scratch with Python
This video from AssemblyAI will be our main guide for the implementation. It breaks down the process into manageable parts: creating the class, fitting the model with bootstrapping, and making predictions with majority voting.
Watch from the beginning to 10:48. You can pause and code along. Focus on the three main components of the implementation: Initialization (__init__) (01:45 - 03:34): Setting up the class with hyperparameters like n_trees, max_depth, etc., and an empty list to hold the trees. Training (fit) (03:34 - 07:17): The core training loop where you'll create bootstrapped data samples and train a decision tree on each. Prediction (predict) (07:17 - 10:48): Aggregating the predictions from all trees and performing a majority vote.
Let's break down the implementation based on the video's logic.
Step 1: The RandomForest Class Structure
First, we define our class and its __init__ method. It will store the hyperparameters that control the forest's behavior and the individual trees within it.
import numpy as np
from collections import Counter
# We assume you have a DecisionTree class from a previous lesson in a file named decision_tree.py
from decision_tree import DecisionTree
class RandomForest:
def __init__(self, n_trees=10, max_depth=10, min_samples_split=2, n_features=None):
self.n_trees = n_trees
self.max_depth = max_depth
self.min_samples_split = min_samples_split
self.n_features = n_features
self.trees = [] # To store the individual trained decision trees
n_trees: The number of decision trees in the forest (theMfrom our theory lesson).max_depth,min_samples_split: These are hyperparameters for the individual decision trees.n_features: The number of features to consider at each split. IfNone, it defaults to all features (like standard bagging).
Step 2: The fit Method and Bootstrapping
The fit method orchestrates the training process. It iterates n_trees times, and in each iteration, it creates a new bootstrapped dataset and trains a new decision tree on it.
def fit(self, X, y):
"""Trains the random forest."""
self.trees = []
for _ in range(self.n_trees):
# 1. Create a Decision Tree instance
tree = DecisionTree(
min_samples_split=self.min_samples_split,
max_depth=self.max_depth,
n_features=self.n_features
)
# 2. Create a bootstrapped sample
X_sample, y_sample = self._bootstrap_samples(X, y)
# 3. Train the tree on the sample
tree.fit(X_sample, y_sample)
# 4. Add the trained tree to our forest
self.trees.append(tree)
def _bootstrap_samples(self, X, y):
"""Creates a bootstrap sample of the dataset."""
n_samples = X.shape[0]
# Sample indices with replacement
idxs = np.random.choice(n_samples, size=n_samples, replace=True)
return X[idxs], y[idxs]
The _bootstrap_samples helper function is the heart of the bagging process. It generates a new dataset of the same size as the original by sampling with replacement.
Step 3: The predict Method and Majority Voting
Once the forest is trained, making a prediction involves getting a vote from each tree and finding the most common one.
def predict(self, X):
"""Makes predictions for a set of samples."""
# Get predictions from all trees
tree_preds = np.array([tree.predict(X) for tree in self.trees])
# tree_preds shape: (n_trees, n_samples)
# We need to get the majority vote for each sample.
# Transpose so shape is (n_samples, n_trees)
tree_preds = tree_preds.T
# For each sample, find the most common prediction (label)
y_pred = [self._most_common_label(tree_pred) for tree_pred in tree_preds]
return np.array(y_pred)
def _most_common_label(self, y):
"""Finds the most frequent label in an array of predictions."""
counter = Counter(y)
most_common = counter.most_common(1)[0][0]
return most_common
The predict logic is straightforward:
- Collect predictions from every tree for all input samples
X. - For each sample, gather all the predictions made for it.
- Use a helper function like
_most_common_labelto determine the majority vote. This becomes the final prediction for that sample.
You can find complete, runnable code examples in the resources below. They offer slightly different implementations, which can be instructive to compare.
Random_Forest_From_Scratch by Eric-D-Stevens
This GitHub repository provides a well-documented, from-scratch implementation of a Random Forest. It's a great reference to see the code structured slightly differently, for example by using a _TreeBag helper class.
Navigate to the section 'Python Implementation of Random Forest'. You don't need to read everything, but focus on the structure of the Forest class, its __init__ method, and the predict function. Notice how it encapsulates the tree and its associated features in a _TreeBag subclass.
Test your understanding!
In our predict method, we transposed the tree_preds array from (n_trees, n_samples) to (n_samples, n_trees). Why is this step crucial for the majority voting logic that follows? What would happen if we tried to iterate and find the most common label without transposing?
Show answer
The transposition is crucial because it groups the predictions by sample. Before transposing, each row represents a single tree's predictions across all samples. After transposing, each row represents all the different tree predictions for a single sample.
If we didn't transpose, iterating through the (n_trees, n_samples) array would mean that _most_common_label would be called on each row. This would calculate the most common prediction made by a single tree across all the test samples, which is not what we want. We need to find the consensus prediction for each sample across all the trees.
3. Evaluation and Out-of-Bag Error
Now that we have a fully implemented RandomForest class, we can train and test it. The final part of the AssemblyAI video demonstrates this.
How to implement Random Forest from scratch with Python
Let's see the code in action. This final segment shows how to use our RandomForest class with a real dataset.
Watch from 10:48 to the end. Observe how the classifier is instantiated, fit on training data, and evaluated on test data. Notice how easily you can tune hyperparameters like n_trees.
A "Free" Validation Set: Out-of-Bag Error
One of the most elegant properties of bagging (and thus Random Forests) is a technique for model evaluation called Out-of-Bag (OOB) error.
Remember how bootstrapping works? On average, each bootstrap sample contains about 63.2% of the original data points. This means that for any given data point in your training set, it was left out of the training process for about 36.8% of the trees. These are its "out-of-bag" trees.
We can use these OOB trees to get an unbiased estimate of the model's performance without needing a separate validation set. The process is:
- For each data point in the training set, identify all the trees that did not use it for training.
- Let those trees make a prediction for that data point.
- Take a majority vote among those predictions to get a single "OOB prediction".
- Compare this OOB prediction to the true label.
- The OOB error is the proportion of incorrect OOB predictions across the entire training set.
This powerful technique gives you a reliable estimate of test-set performance "for free" during the training process.
Machine Learning Lecture 31 "Random Forests / Bagging" -Cornell CS4780 SP17
For a deeper dive into the theory and power of OOB error, let's return to the lecture from Kilian Weinberger.
Watch the segment from 21:25 to 28:07. This is a more theoretical explanation, but it clearly articulates why OOB error is an unbiased estimate of the test error and why this is such a powerful advantage of bagging-based methods.
Conclusion
Congratulations! You have now not only understood the theory behind Random Forests but have also walked through a complete from-scratch implementation. This solidifies your grasp of ensemble methods and gives you a powerful tool for your machine learning arsenal.
Key Takeaways:
- A Random Forest is an ensemble of Decision Trees that uses bagging and feature subsampling to build a robust classifier.
- Bagging (bootstrapping data) provides the primary mechanism for creating diverse models.
- Feature subsampling at each split is the crucial step that decorrelates the trees, making the ensemble more effective than simple bagging.
- The implementation involves a main class that manages a list of tree objects, a
fitmethod to handle the bootstrapping and training loop, and apredictmethod that aggregates votes. - Out-of-Bag (OOB) error is a powerful feature of bagged models, providing a reliable estimate of test performance without a separate validation set.
Preview of the next lesson:
We've now seen how to combine models in parallel with bagging to reduce variance. In the next lesson, we will shift our focus to the other major ensemble strategy: boosting. We will implement AdaBoost and Gradient Boosting Machines (GBMs) from first principles to understand how training models sequentially can powerfully reduce bias and create some of the highest-performing models in classical machine learning.