Skip to main content
Create your own

Implementing Logistic Regression for Binary Classification

Hello! Let's continue our journey into classical machine learning algorithms.

Introduction

In our previous lesson, we built the linear regression model from the ground up, learning how to predict continuous values. We also tackled the problem of overfitting using L1 (Lasso) and L2 (Ridge) regularization. You are now proficient with the concepts of a linear model, a cost function (MSE), and optimization via gradient descent.

But what if our goal isn't to predict a house price, but to classify an email as spam or not spam? This is a binary classification problem, where the target is a discrete category (e.g., 0 or 1), not a continuous number.

This lesson directly addresses the learning outcome: Implement logistic regression for binary classification tasks. We will see how the linear model you just mastered can be cleverly adapted for this new challenge.

We will cover:

  1. From Linear to Logistic: How to transform the unbounded output of a linear model into a valid probability using the sigmoid function.
  2. The Right Cost Function: Why Mean Squared Error is a poor choice for classification and how to derive the cross-entropy loss from the principle of maximum likelihood.
  3. Gradient Descent for Classification: We'll compute the gradient of the cross-entropy loss and see a surprising and elegant similarity to the gradient for linear regression.
  4. Implementation from Scratch: We will translate the theory into a working Python implementation using NumPy.

This lesson builds directly on your understanding of linear models and gradient descent, forming a bridge from regression to classification.

1. From Linear Prediction to Probability: The Sigmoid Function

Our linear regression model produces a prediction . The output, let's call it , can be any real number from to . For binary classification, we need to map this output to a probability that the sample belongs to class 1, which must be in the range [0, 1].

This is where the logistic function, more commonly known in machine learning as the sigmoid function, comes in. It takes any real-valued number and "squashes" it into the (0, 1) range.

Comparison of Linear and Logistic Regression Models
This image visually contrasts the output of linear regression, which is a continuous line, with logistic regression, which produces an S-shaped curve bounded between 0 and 1. This "S-curve" is the sigmoid function.

The sigmoid function, denoted by , is defined as:

The hypothesis of our logistic regression model is then the sigmoid function applied to the linear part:

This is our model's predicted probability that the true label is 1. To make a final classification, we apply a threshold, typically 0.5.

  • If , we predict class 1.
  • If , we predict class 0.

The following video provides an excellent introduction to this core idea.

Logistic Regression From Scratch in Python (Mathematical)

Let's watch a segment from the video 'Logistic Regression From Scratch in Python' by NeuralNine. It explains why logistic regression is used for classification despite its name and introduces the sigmoid function as the key component that transforms a linear output into a probability.

Watch from 02:54 to 14:11. This covers the motivation for logistic regression, the linear model at its core, and the role and mathematics of the sigmoid function. Focus on understanding how the linear combination of features, just like in linear regression, is fed into the sigmoid function to produce a probability.

This thresholding naturally defines a decision boundary. Notice that whenever . Therefore, our model predicts class 1 when . The equation defines a line (or a hyperplane in higher dimensions) that separates the feature space into two regions, one for each class.

2. A Cost Function for Classification: Cross-Entropy Loss

Now that we have a prediction, we need a cost function to tell us how "wrong" it is. Our old friend, Mean Squared Error (MSE), is not a good fit here. If we were to use MSE with our sigmoid-based predictions, the resulting cost function would be non-convex, meaning it would have many local minima. This would make it very difficult for gradient descent to find the global minimum.

Instead, we need a cost function that is convex and specifically designed for probability-based predictions. This function is the Binary Cross-Entropy Loss (or Log Loss), and it can be derived from the principle of Maximum Likelihood Estimation (MLE).

The goal of MLE is to find the model parameters () that maximize the likelihood (probability) of observing our given training data. For binary classification, this leads to the cross-entropy cost function, which for a single training example is:

Let's analyze this:

  • If the true label : The loss becomes . If our prediction is close to 1 (correct), is close to 0. If is close to 0 (very wrong), approaches infinity.
  • If the true label : The loss becomes . If our prediction is close to 0 (correct), is close to 1, and the loss is near 0. If is close to 1 (very wrong), is close to 0, and the loss approaches infinity.

This function perfectly penalizes confident but incorrect predictions. The total cost function is the average loss over all training examples:

The following video segment provides a detailed derivation of this function.

Logistic Regression From Scratch in Python (Mathematical)

Let's return to the NeuralNine video to see how this cross-entropy loss function is derived from the likelihood function.

Watch from 14:11 to 24:10. The video walks through constructing the likelihood function, applying the logarithm to simplify it, and negating it to create a cost function for minimization. Follow the mathematical steps closely.

Test your understanding!

A logistic regression model is trying to predict if an image contains a cat (y=1) or not (y=0).

  1. For an image of a cat, the model predicts a probability of 0.1.
  2. For an image of a dog, the model predicts a probability of 0.9.

Which prediction results in a higher loss? Calculate the loss for each case using the formula Loss = -[y log(Å·) + (1-y) log(1-Å·)]. Use the natural logarithm (np.log in Python).

Show answer
  1. Case 1 (Cat image): True label y=1, prediction Å·=0.1.
    The loss is - [1 * log(0.1) + (0) * log(0.9)] = -log(0.1) ≈ 2.30.

  2. Case 2 (Dog image): True label y=0, prediction Å·=0.9.
    The loss is - [0 * log(0.9) + (1) * log(1-0.9)] = -log(0.1) ≈ 2.30.

Both predictions result in the same high loss. This demonstrates the symmetry of the cross-entropy function: being confidently wrong is penalized heavily, regardless of the class.

3. Gradient Descent with Cross-Entropy Loss

We have our cost function. Now, to use gradient descent, we need its partial derivatives with respect to our parameters and . This involves applying the chain rule, just as you saw in calculus. The derivation is a bit involved, but it leads to a remarkably simple and elegant result.

We want to compute . Using the chain rule, this can be broken down:

When you work through the math, many terms cancel out beautifully.

Logistic Regression From Scratch in Python (Mathematical)

Let's watch the final mathematical segment from the NeuralNine video. It shows the derivation of the gradient for the logistic regression cost function.

Watch from 24:10 to 35:15. This part can be mathematically dense, but the key is to follow the application of the chain rule. Focus on the final, simplified result for the gradient.

After all the calculus, the partial derivative with respect to a single weight is:

And the partial derivative with respect to the bias is:

This is a powerful result. The gradient calculation looks identical to the one for linear regression with MSE! The only difference is how is defined. In linear regression, it was ; here, it is . This deep connection showcases the elegance of generalized linear models.

4. Implementation from Scratch

Now we have all the components to build our logistic regression model in Python using NumPy. We'll implement a class that encapsulates the model's logic.

The core components will be:

  • An __init__ method to store hyperparameters like the learning rate and number of iterations.
  • A _sigmoid helper function.
  • A fit method that implements gradient descent:
    • Initializes weights and bias.
    • Loops for a set number of iterations.
    • Inside the loop, it calculates predictions, computes the gradients using the formula above, and updates the weights and bias.
  • A predict method that takes new data, calculates probabilities using the trained weights, and returns binary predictions based on the 0.5 threshold.

Let's watch a full implementation that puts these pieces together.

Logistic Regression From Scratch in Python (Mathematical)

The 'Logistic Regression From Scratch in Python' video by NeuralNine provides an excellent, clean, and vectorized implementation. We'll use it as our guide.

Watch from 35:15 to 48:00. This section translates all the math we've discussed into Python code. Pay close attention to: Vectorization: How the gradient calculation is performed efficiently on the entire dataset using NumPy's matrix operations (@ operator for matrix multiplication, .T for transpose). This is a direct parallel to your experience with optimized computations in CS. Bias Handling: The video adds a bias column (a column of ones) to the feature matrix X. This is a common trick that allows the bias b to be treated as just another weight w_0, simplifying the update equations. Training and Evaluation: The code trains the model on a real dataset from scikit-learn and evaluates its accuracy, showing that our from-scratch implementation works effectively.

Here's a simplified functional version of the core gradient update logic, reflecting what you saw in the video:

import numpy as np

def sigmoid(z):
    return 1 / (1 + np.exp(-z))

# Assume X_train (m, n), y_train (m,), learning_rate, and epochs are defined

# 1. Initialize parameters
m, n = X_train.shape
weights = np.zeros(n)
bias = 0

# 2. Gradient Descent Loop
for _ in range(epochs):
    # Calculate linear model and apply sigmoid
    z = np.dot(X_train, weights) + bias
    y_pred = sigmoid(z) # Shape (m,)

    # Calculate gradients
    dw = (1 / m) * np.dot(X_train.T, (y_pred - y_train)) # Shape (n,)
    db = (1 / m) * np.sum(y_pred - y_train)

    # Update parameters
    weights -= learning_rate * dw
    bias -= learning_rate * db

print("Training complete!")
print("Final weights:", weights)
print("Final bias:", bias)

This snippet captures the essence of the fit method. The full implementation in the video, organized into a class, is a more robust and reusable approach.

Conclusion

Congratulations! You have successfully extended your knowledge from linear regression to build a powerful classification algorithm.

Key Takeaways:

  • Logistic Regression is a fundamental algorithm for binary classification that models the probability of an outcome.
  • It combines a linear predictor () with the sigmoid function to produce a probability .
  • The appropriate cost function is the Binary Cross-Entropy Loss, which is derived from maximum likelihood principles and is convex for this problem.
  • Optimization is performed using gradient descent, and the gradient formula is nearly identical in form to that of linear regression, highlighting a deep connection between the models.

Preview of the next lesson:
So far, we have looked at discriminative models like logistic regression, which learn a direct mapping from inputs to outputs by finding a decision boundary. In our next lesson, we will explore a different approach with Naive Bayes classifiers. Naive Bayes is a generative model, meaning it learns the underlying probability distribution of the data for each class. This probabilistic approach offers a new perspective on classification.

Can't find a good explanation? Sign up and we'll make it for you

Sign up