Skip to main content
Create your own

Eigenvalues and Eigenvectors: ML Significance

Hello! Welcome to the second lesson in our journey through the mathematical foundations of AI.

In our last session, we explored how matrices act as linear transformations, stretching, squashing, and rotating vector spaces. We concluded by posing a question: are there any vectors that are special, in that a transformation only scales them without changing their direction?

Today, we'll answer that question. This lesson is all about eigenvectors and eigenvalues. You will learn how to calculate these special vectors and their corresponding scaling factors, and, more importantly, you will understand their profound significance in machine learning. They are key to unlocking the underlying structure of data and are the foundation for powerful techniques like Principal Component Analysis (PCA).

What are Eigenvectors and Eigenvalues?

An eigenvector of a matrix is a non-zero vector whose direction remains unchanged when the linear transformation represented by the matrix is applied to it. It may be stretched, squished, or flipped, but it won't be rotated off its original line (its "span"). The factor by which the eigenvector is scaled is called the eigenvalue.

This relationship is elegantly captured by the defining equation:

The fundamental relationship between a transformation matrix \(A\), an eigenvector \(\mathbf{v}\), and its corresponding eigenvalue \(\lambda\). The transformation \(A\) only scales the eigenvector \(\mathbf{v}\) by the factor \(\lambda\).

To build a strong intuition for this concept, let's watch one of the best explanations available, from 3Blue1Brown.

Eigenvectors and eigenvalues | Chapter 14, Essence of linear algebra

This video, 'Eigenvectors and eigenvalues', provides a fantastic visual and conceptual introduction. The first part will show you what eigenvectors and eigenvalues represent geometrically.

Starting about a minute in, watch the introduction to eigenvectors. Focus on how certain vectors stay on their own span after a transformation, while most are knocked off.

As you saw, eigenvectors are the "axes" of a linear transformation. They reveal the directions along which the transformation acts purely by stretching or compressing.

So why is finding these special axes useful? In physics, if you consider a 3D rotation, the eigenvector is the axis of rotation—a concept much more intuitive than the full 3x3 rotation matrix. In machine learning, they help us find the most important directions in our data.

The Calculation of Eigenvalues and Eigenvectors

Now that we have the conceptual understanding, let's move to the mechanics of how to find them. The entire process starts from the defining equation, .

Our goal is to find the values for (the eigenvalues) and (the eigenvectors) that satisfy this equation for a given square matrix .

We can rearrange the equation as follows:

  1. To factor out , we need to express as a matrix-vector product. We use the identity matrix :
  2. Now, we can factor out :

This equation tells us that the matrix transforms the vector into the zero vector. We are looking for a non-zero eigenvector , which is a "non-trivial" solution. From our knowledge of linear algebra, a matrix equation has a non-zero solution for only if the matrix is "singular" or "non-invertible". This happens precisely when the determinant of the matrix is zero.

Therefore, to find our eigenvalues, we need to find the values of that make the determinant of equal to zero:

This is known as the characteristic equation. The following video segment walks through this derivation and shows how it leads to finding the eigenvalues.

Eigenvectors and eigenvalues | Chapter 14, Essence of linear algebra

Let's return to the 3Blue1Brown video to see how the characteristic equation is derived and used.

Starting around five minutes in, watch the computational overview. Pay close attention to how setting the determinant to zero is the key step to ensure we can find a non-zero eigenvector.

Step-by-Step Guide

Let's solidify this with a step-by-step process. The following video from Professor Dave Explains provides a very clear, practical walkthrough of the calculations for a 2x2 matrix.

Finding Eigenvalues and Eigenvectors

This video, 'Finding Eigenvalues and Eigenvectors', will guide you through the computational steps.

Watch from about three minutes in. The first part covers finding eigenvalues by solving the characteristic polynomial. The second part shows how to substitute those eigenvalues back into the equation to find corresponding eigenvectors.

To summarize the process:

  1. Find the Eigenvalues ():

    • Set up the characteristic equation: .
    • Solving this equation will yield a polynomial in . The roots of this polynomial are the eigenvalues of matrix .
  2. Find the Eigenvectors ():

    • For each eigenvalue you found, substitute it back into the equation .
    • This will give you a system of linear equations. Solve this system to find the vector .
    • Note that any non-zero scalar multiple of an eigenvector is also an eigenvector. So, there isn't a single unique eigenvector for a given eigenvalue, but rather an entire line (or subspace) of them. We typically pick a simple representative vector.
Test your understanding!

Let's calculate the eigenvalues and eigenvectors for the matrix .

  1. Find the eigenvalues.
  2. Find the corresponding eigenvectors.
Show answer
  1. Find Eigenvalues:
    First, we set up the characteristic equation .

    Factoring the quadratic equation:

    So, the eigenvalues are and .

  2. Find Eigenvectors:

    • For :
      We solve .

      Both rows give the same equation: (or ), which means . A simple choice for the eigenvector is .

    • For :
      We solve .

      This gives the equation , which means . A simple choice for the eigenvector is .

Significance in Machine Learning

This might seem like a purely mathematical exercise, but it has profound implications in data science. In ML, we often aren't interested in an arbitrary matrix , but in a special one: the covariance matrix.

For a given dataset, the covariance matrix is a square matrix that describes the variance of each feature and the covariance between pairs of features.

  • The eigenvectors of the covariance matrix point in the directions of maximum variance in the data. These are called the principal components.
  • The eigenvalue corresponding to each eigenvector tells you the amount of variance captured by that direction. A large eigenvalue means that the data is very spread out along that eigenvector.

This is the core idea behind Principal Component Analysis (PCA), one of the most widely used techniques for dimensionality reduction.

An illustration of eigenvectors on a dataset. Eigenvector 1 (red) corresponds to the largest eigenvalue and points along the direction of maximum variance. Eigenvector 2 (green) is orthogonal and captures the next largest variance. PCA uses these to find the most important "directions" in the data.

By finding the eigenvectors and eigenvalues of the covariance matrix, we can identify the directions that contain the most information (variance) and discard the ones that contain the least, reducing the number of dimensions while preserving the essential structure of the data.

Eigenvectors and Eigenvalues: Key Insights for Data Science

The article 'Eigenvectors and Eigenvalues: Key Insights for Data Science' discusses several applications of these concepts in ML.

Read the section titled 'Considerations in Data Science and Machine Learning'. Read through the practical applications—it provides a concise list of how eigenvectors and eigenvalues are used in various fields, reinforcing their practical importance.

Implementation in Python

Given your background in software engineering, you know that we rarely perform these calculations by hand. Scientific computing libraries like NumPy make this trivial. Let's see how to compute the eigenvalues and eigenvectors for the matrix from our exercise using Python.

Eigenvectors and Eigenvalues - Detailed Explanation on ...

This article from MachineLearningPlus shows the practical implementation using Python's NumPy library.

Look at the Python code block in the middle of the page. You can try running the Python implementation yourself. It uses np.linalg.eig to find the eigenvalues and eigenvectors of a matrix.

Here is the code for our example matrix :

import numpy as np




# Define the matrix
A = np.array([[4, -2],
              [1,  1]])




# Calculate eigenvalues and eigenvectors
eigenvalues, eigenvectors = np.linalg.eig(A)

print("Eigenvalues:", eigenvalues)
print("Eigenvectors:")
print(eigenvectors)




# Note: NumPy returns eigenvectors as columns of the matrix.
# The first column is the eigenvector for the first eigenvalue.
# Let's verify for the first eigenvalue (3) and its eigenvector
lambda_1 = eigenvalues[0]
v_1 = eigenvectors[:, 0]
print("\nVerifying A*v_1 = lambda_1*v_1:")
print("A @ v_1:", A @ v_1)
print("lambda_1 * v_1:", lambda_1 * v_1)

The output confirms that our manual calculations were correct (up to a scaling factor and normalization for the eigenvectors).

Eigenbasis and Diagonalization

Finally, let's connect back to the idea of transformations and prepare for our next lesson. What happens if a matrix has enough eigenvectors to form a basis for the entire space? Such a basis is called an eigenbasis.

If we change our coordinate system to use this eigenbasis, the transformation becomes incredibly simple. In the eigenbasis, the transformation matrix is a diagonal matrix with the eigenvalues on the diagonal. All it does is scale the new basis vectors.

This process of simplifying a matrix by changing to its eigenbasis is fundamental to eigendecomposition, a topic we'll explore next.

Eigenvectors and eigenvalues | Chapter 14, Essence of linear algebra

The final part of the 3Blue1Brown video explains this powerful idea of an eigenbasis.

Starting at 12:44, watch the eigenbasis explanation. This section demonstrates how changing to an eigenbasis diagonalizes the transformation matrix, making complex operations like repeated matrix multiplication much easier.

Not all matrices can be diagonalized (e.g., a shear transformation doesn't have enough eigenvectors to span the space), but for many matrices we encounter in machine learning (like symmetric covariance matrices), this is possible and extremely useful.

Conclusion

In this lesson, we've gone from the intuitive idea of "special vectors" to the practical calculations and their significance in machine learning.

Key Takeaways:

  • An eigenvector of a matrix is a vector that satisfies , where is its corresponding eigenvalue. It represents a direction that is only scaled by the transformation.
  • Eigenvalues are found by solving the characteristic equation, .
  • Eigenvectors are found by substituting each eigenvalue back into and solving for .
  • In machine learning, the eigenvectors of a covariance matrix are the principal components of the data, representing directions of maximum variance. The eigenvalues quantify this variance.
  • An eigenbasis (a basis of eigenvectors) simplifies a linear transformation into a diagonal matrix of eigenvalues, making computations easier.

Preview of the next lesson:

We've just seen that we can simplify a matrix by changing to its eigenbasis. In the next lesson, we will formalize this into a powerful technique called eigendecomposition. We will then explore an even more general and widely applicable method, Singular Value Decomposition (SVD), which can decompose any matrix, not just square ones.

Can't find a good explanation? Sign up and we'll make it for you

Sign up