Skip to main content
Create your own

Regularized Linear Regression

Hello! Let's dive into the first lesson of Module 3.

Introduction

In our previous lesson, we mastered the tools for evaluating regression models, focusing on metrics like Mean Squared Error (MSE), Mean Absolute Error (MAE), and R-squared. We now know how to measure a model's performance, but how do we build one in the first place?

Today, we'll build the most fundamental regression algorithm: Linear Regression. This lesson addresses the learning outcome: Implement linear regression and apply L1/L2 regularization to prevent overfitting.

We will cover:

  1. Implementing Linear Regression: We'll define the model and explore two methods for finding the optimal parameters: the analytical "Normal Equation" and the iterative "Gradient Descent" algorithm.
  2. The Problem of Overfitting: We'll see how a linear model can become too complex and fail to generalize.
  3. L2 Regularization (Ridge Regression): You'll learn to add a penalty term to shrink model coefficients and prevent overfitting.
  4. L1 Regularization (Lasso Regression): We'll explore a different penalty that can shrink coefficients all the way to zero, effectively performing feature selection.

This lesson connects directly to our previous work. The MSE you learned to calculate is the very cost function we will now learn to minimize. Your background in computer science and mathematics will be particularly useful here, as we'll touch on linear algebra, calculus, and optimization algorithms.

1. Linear Regression: The Baseline Model

Linear regression assumes a linear relationship between input features () and a continuous target variable (). For a single data point with features, the model's prediction, , is:

Using vector notation, which you'll be familiar with from Module 1, this simplifies to:

Here, is the vector of weights, and is the bias (or intercept). Our goal is to find the optimal values for and that make our predictions as close as possible to the actual values across all training data. To measure this "closeness," we use the Mean Squared Error (MSE) as our cost function, .

Finding the parameters that minimize this cost function is the core task of training. There are two primary ways to do this.

Method 1: The Analytical Solution (Normal Equation)

For linear regression, we are fortunate to have a direct, analytical solution. By setting the gradient of the cost function with respect to the weights to zero () and solving for , we can derive a closed-form solution called the Normal Equation.

First, we augment our feature matrix to include a column of ones for the bias term. This allows us to combine and into a single weight vector. The equation is:

  • Pros: It's an exact, one-step calculation. No need to choose a learning rate or iterate.
  • Cons: The matrix inversion operation is computationally expensive, typically with a complexity of , where is the number of features. This makes it impractical for datasets with many thousands of features. It also fails if the matrix is singular (non-invertible), which can happen if you have redundant features (i.e., perfect multicollinearity).

Method 2: The Iterative Solution (Gradient Descent)

A more general and scalable approach is Gradient Descent, which you implemented in the previous module. We start with random weights and iteratively update them by taking small steps in the direction of the negative gradient of the cost function.

The update rule for each weight is:

Where is the learning rate.

The partial derivative of the MSE cost function with respect to a weight is:

This iterative process slowly descends the cost function surface to find the minimum, as visualized below.

Path of Converging Gradient Descent on a Cost Function Surface
This 3D plot illustrates the path of gradient descent. The algorithm starts at a high-cost point (red area) and iteratively updates the weights (w0, w1) to move towards the minimum of the cost function (blue area).

Gradient descent is the workhorse of modern machine learning, as it scales to millions of features and is the basis for training complex neural networks.

2. Overfitting and the Need for Regularization

Simple linear regression works well when the number of features is small and they are genuinely predictive. However, if we use too many features, or create complex polynomial features (e.g., ), our model can become overly complex. It might learn the noise in the training data instead of the underlying signal. This is overfitting.

An overfit model will have a very low training error (low MSE) but perform poorly on new, unseen data (high validation/test error). This happens because the model's coefficients () grow to very large positive or negative values to perfectly fit every training point.

Regularization is a technique to prevent this. It works by adding a penalty term to the cost function that discourages large coefficient values.

6. L1 & L2 Regularization

To start our discussion on regularization, watch this introductory clip from the video 'L1 & L2 Regularization' from Inside Bloomberg. It frames regularization as a method to intentionally fit the training data less well in order to achieve better generalization.

Watch the first 2 minutes and 12 seconds of the video. This sets the stage for everything that follows by explaining the core motivation behind regularization.

The general form of a regularized cost function is:

Here, (lambda) is a hyperparameter that controls the strength of the regularization. A higher imposes a stronger penalty. The two most common ways to define the complexity of are the L2 norm and the L1 norm.

3. L2 Regularization (Ridge Regression)

Ridge Regression uses the squared magnitude of the coefficients as its penalty term. This is the L2 norm of the weight vector.

The Ridge cost function is:

(Note: The bias term is typically not regularized).

By adding this penalty, we force the optimization process to find a balance between fitting the data (minimizing MSE) and keeping the weights small (minimizing the L2 norm).

Analytical Solution for Ridge

One of the elegant properties of Ridge is that it still has a closed-form solution. The penalty term ensures the matrix is always invertible.

Understanding L1 and L2 regularization with analytical and ...

The article 'Understanding L1 and L2 regularization...' provides a clean mathematical derivation for the Ridge Regression solution. Let's look at how the normal equation is modified.

Read the short section '3.1 Analytical derivation of L2 regularization'. Focus on the final formula, which shows how the identity matrix and lambda are introduced.

As derived in the article, the closed-form solution for Ridge regression is:

Notice the addition of (where is the identity matrix) inside the parentheses. This small addition makes the matrix invertible and is the mathematical core of Ridge regression.

Implementing Ridge Regression

Now, let's turn theory into practice. The following video is a tutorial on implementing Ridge from scratch in Python. It's an excellent guide that covers not just the formula but also crucial practical steps like data standardization.

Ridge Regression From Scratch In Python [Machine Learning Tutorial]

The video 'Ridge Regression From Scratch In Python' by Dataquest will walk you through a full implementation. We'll focus on the core logic.

Watch from 07:04 to 16:44. This segment is the heart of the implementation. Pay close attention to: Why data standardization is necessary before applying regularization (07:04-09:00, 14:15-16:00). How the Ridge penalty (lambda * I) is added to the X transpose X matrix in the code (12:08-13:30, 21:55-22:45).

Here's a summary of the implementation steps in Python using NumPy:

  1. Standardize Data: Subtract the mean and divide by the standard deviation for each feature column. This is critical because regularization penalizes coefficients indiscriminately. If features are on different scales (e.g., age in years and income in dollars), the coefficient for income will naturally be smaller, and regularization would unfairly penalize the feature with the larger-scale units.
  2. Add Intercept: Add a column of ones to your feature matrix to account for the intercept term.
  3. Construct Penalty Matrix: Create an identity matrix of size (where is the number of features), set the top-left element to zero (to not penalize the intercept), and multiply by .
  4. Solve: Implement the formula using np.linalg.inv() for the inverse and the @ operator for matrix multiplication.
import numpy as np

# Assume X_train_scaled, y_train are pre-processed
# X_train_scaled has shape (n_samples, n_features)
# y_train has shape (n_samples,)

def fit_ridge(X, y, alpha):
    # 1. Add intercept term to X
    X_b = np.c_[np.ones((X.shape[0], 1)), X]
    d = X_b.shape[1] # Number of features + intercept

    # 2. Create penalty matrix
    I = np.identity(d)
    I[0, 0] = 0 # Don't regularize the intercept term

    # 3. Apply the Normal Equation with L2 penalty
    # (X.T @ X + alpha * I)^-1 @ X.T @ y
    A = X_b.T @ X_b + alpha * I
    w = np.linalg.inv(A) @ X_b.T @ y
    
    return w

# Example usage
alpha = 1.0 # Regularization strength
weights = fit_ridge(X_train_scaled, y_train, alpha)
print("Ridge Weights:", weights)

4. L1 Regularization (Lasso Regression)

Lasso Regression (Least Absolute Shrinkage and Selection Operator) uses the L1 norm of the weight vector as its penalty.

The Lasso cost function is:

The key difference from Ridge is the use of the absolute value instead of the square . This seemingly small change has a profound effect: Lasso can shrink some coefficients to exactly zero. This means it automatically performs feature selection, creating a sparse model.

The Geometric Intuition Behind Sparsity

Why does L1 lead to sparsity while L2 doesn't? The answer lies in the geometry of the constraint regions.

Geometric Comparison of L1 and L2 Regularization
This image visualizes the difference between L2 and L1 regularization. The ellipses are the contours of the unregularized MSE cost function. The constraint region for L2 is a circle, while for L1 it's a diamond. The optimal solution is the first point where the ellipses touch the constraint region. For L1, this is very likely to happen at a corner, where one of the weights is zero, resulting in a sparse solution.

As the image shows, the circular L2 constraint is likely to touch the elliptical cost function contours at a point where neither weight is zero. In contrast, the sharp corners of the L1 diamond constraint lie on the axes. It is much more probable that the contours will first make contact with the constraint region at one of these corners, forcing the corresponding weight to be zero.

6. L1 & L2 Regularization

Let's revisit the 'L1 & L2 Regularization' video for a more detailed walkthrough of this geometric argument and to see why sparsity is desirable.

Watch from 23:58 to 33:47 to see the introduction to Lasso and the benefits of sparsity. Then, continue from 33:47 to 49:32 for the classic geometric explanation of why Lasso produces sparse solutions.

The primary benefits of a sparse model from Lasso are:

  • Feature Selection: It automatically identifies and removes irrelevant features.
  • Interpretability: Models with fewer features are easier to understand and explain.
Test your understanding!

You are building a model to predict house prices using 500 features, many of which you suspect are redundant or useless (e.g., "distance to nearest park" and "distance to nearest green space"). Would you prefer Ridge or Lasso, and why?

Show answer

You would likely prefer Lasso. With a large number of features, many of which might be irrelevant, Lasso's ability to drive coefficients to exactly zero acts as a form of automatic feature selection. This would result in a simpler, more interpretable model that only uses the most important predictors. Ridge, in contrast, would shrink the coefficients of the useless features towards zero but would never make them exactly zero, keeping all 500 features in the model.

A Probabilistic Perspective

Another deep insight comes from a Bayesian statistical view. As you may recall from Module 1, we can use prior beliefs about our parameters.

  • Ridge (L2) is equivalent to Maximum A Posteriori (MAP) estimation assuming a Gaussian prior on the weights. The Gaussian distribution is smooth and prefers weights to be near zero but doesn't strongly enforce it.
  • Lasso (L1) is equivalent to MAP estimation assuming a Laplace prior on the weights. The Laplace distribution has a sharp peak at zero, which provides a much stronger "push" for weights to be exactly zero.

You can read more about this in sections 2.2 and 3.2 of the Understanding L1 and L2 regularization... article if you're curious to see the mathematical derivation.

Optimization Challenge with Lasso

Lasso presents a new challenge: the L1 norm, , is not differentiable at . This means we can't use standard gradient descent because the gradient is undefined.

The most common solution is an algorithm called Coordinate Descent.

6. L1 & L2 Regularization

Let's watch one last clip from the 'L1 & L2 Regularization' video to understand how coordinate descent solves the Lasso optimization problem.

Watch from 1:12:25 to 1:21:00. This section introduces coordinate descent and the 'shooting algorithm' for Lasso. The key takeaway is that we can optimize one coefficient at a time, and for each single-coefficient problem, there is a simple closed-form solution (called a soft-thresholding operator).

In coordinate descent, we iterate through each coefficient, one at a time, and update it to the value that minimizes the cost function while keeping all other coefficients fixed. The magic of Lasso is that this one-variable optimization problem has a simple, closed-form solution. We repeat this cycle until the coefficients converge. While implementing coordinate descent is more involved, understanding the principle is key to knowing how Lasso models are trained in practice.

Conclusion

In this lesson, you've moved from evaluating models to building one. We implemented the foundational linear regression model and equipped it with powerful regularization techniques to combat overfitting.

Key Takeaways:

  • Linear Regression can be optimized using the analytical Normal Equation or iteratively with Gradient Descent.
  • Regularization adds a penalty for model complexity to the cost function, preventing overfitting.
  • L2 Regularization (Ridge) uses a penalty. It shrinks coefficients towards zero, improving generalization, especially when features are correlated.
  • L1 Regularization (Lasso) uses a penalty. It can shrink coefficients to exactly zero, performing automatic feature selection and creating sparse, interpretable models.
  • The choice between Ridge and Lasso depends on your goal: Ridge is a good default, while Lasso is excellent when you have many features and suspect some are irrelevant.

Preview of the next lesson:
Now that we've established a solid foundation with linear regression, we'll see how these ideas can be adapted for classification problems. In the next lesson, you will implement logistic regression, a model that uses a linear equation as its core but maps the output to a probability, allowing it to perform binary classification.

Can't find a good explanation? Sign up and we'll make it for you

Sign up