Skip to main content
Create your own

Object Detection & Segmentation Metrics: mAP & IoU

Hello! Let's dive into our next lesson on evaluating computer vision models.

Introduction

In our last lesson, we built a Mask R-CNN model to perform instance segmentation, generating both bounding boxes and pixel-perfect masks for objects in an image. Building the model is only half the battle. How do we objectively measure its performance? How can we tell if one model is better than another, or if our fine-tuning has actually improved anything?

This lesson addresses exactly that. Your learning outcome is to evaluate object detection and segmentation models using appropriate metrics (mAP, IoU). We will deconstruct the industry-standard metric, mean Average Precision (mAP), from the ground up. This metric is the primary benchmark used in academic papers and competitions like COCO.

We will cover the following concepts in order:

  1. Intersection over Union (IoU): The fundamental measure of overlap that determines if a single prediction is "correct."
  2. Precision and Recall: How these classic classification metrics are adapted for object detection.
  3. Average Precision (AP): A single-number score that summarizes a model's performance for a single class.
  4. Mean Average Precision (mAP): The final metric, which averages AP across all classes and multiple IoU thresholds to give a comprehensive evaluation of the model.

1. Intersection over Union (IoU): The Atomic Unit of Evaluation

Everything starts with a simple question: for a given object, how well does our model's predicted bounding box (or mask) line up with the ground truth? This is measured by Intersection over Union (IoU).

As the name suggests, IoU is calculated as the ratio of the area of overlap between the predicted and ground truth boxes to the area of their combined union.

Visualizing Intersection over Union (IoU)
This image visually demonstrates IoU. A value of 1.0 means a perfect match, while a value of 0 means no overlap at all. The higher the IoU, the better the prediction's localization.

IoU is the foundational metric that allows us to classify a detection as correct or incorrect. We do this by setting an IoU threshold. For example, a common threshold is 0.5. If a predicted box has an IoU with a ground truth box that is greater than or equal to 0.5, we consider it a True Positive (TP). If the IoU is less than 0.5, it's a False Positive (FP).

The same logic applies to segmentation masks. Instead of the area of bounding boxes, we calculate the overlap in terms of pixels.

Intersection Over Union (IoU): From Theory to Practice

The article 'Intersection Over Union (IoU): From Theory to Practice' from Lightly.ai provides an excellent, detailed explanation of IoU.

Please read the 'TL;DR' section and the section titled 'How to Calculate IoU: Step-by-Step'. Focus on: The definition and formula for IoU. The step-by-step calculation for bounding box coordinates. The short Python code example that implements the logic.

2. Precision and Recall in Object Detection

With IoU allowing us to count True Positives and False Positives, we can now define Precision and Recall for object detection.

  • Precision: Of all the detections our model made, what fraction was actually correct?
  • Recall: Of all the actual objects present in the image, what fraction did our model successfully detect? Here, a False Negative (FN) is a ground truth object that the model failed to detect.
Visual Explanation of Precision, Recall, and IoU in Object Detection
This diagram illustrates how Precision, Recall, and IoU are calculated from the predicted 'Detected box' and the 'Object' ground truth. It helps to visualize the ratios involved.

There is a natural trade-off. If we lower our model's confidence threshold, it will produce many more bounding boxes. This will likely increase Recall (we'll miss fewer objects) but decrease Precision (many of the new boxes will be incorrect). The mAP metric is designed to evaluate a model's performance across this entire trade-off.

3. The Path to mAP: Precision-Recall Curve and Average Precision (AP)

To get a single, robust metric, we need to evaluate the model not at a single confidence threshold, but across all possible thresholds. This is done by calculating the Average Precision (AP) for each class.

The calculation of AP is an algorithmic process. For a single class (e.g., "person"):

  1. Gather all predicted bounding boxes for that class from every image in your test set. Each prediction has a confidence score.
  2. Sort these predictions in descending order of their confidence scores.
  3. Iterate down this sorted list. At each prediction, calculate the cumulative Precision and Recall.
  4. Plot these Precision vs. Recall values. This creates the Precision-Recall Curve.
  5. Average Precision (AP) is the area under this Precision-Recall curve.

A higher AP means the model can maintain high precision even as it achieves high recall, which indicates a high-quality model.

Mean Average Precision (mAP) Explained and PyTorch Implementation

The video 'Mean Average Precision (mAP) Explained' by Aladdin Persson provides one of the clearest explanations of this entire process. We'll watch the first part, which walks through a concrete example.

Watch from 00:43 to 07:17. This is a crucial segment. Pay close attention to: (00:43 - 03:13): How predictions are classified as TP or FP based on IoU and then sorted by confidence score. (03:13 - 05:20): The definitions of Precision and Recall in this context. (05:20 - 07:17): The step-by-step calculation of the Precision-Recall curve and how AP is the area under it.

Test your understanding!

In the AP calculation process, why is it necessary to sort the detections by their confidence score before computing the precision-recall curve?

Show answer

Sorting by confidence score simulates moving a decision threshold from high to low. High-confidence predictions are considered first. This allows the precision-recall curve to show how the model's performance changes as we become more lenient and start including lower-confidence detections. A good model will have its true positives ranked with high confidence, leading to high precision at the beginning (low recall) of the curve. Without sorting, the curve would be erratic and wouldn't represent the model's precision/recall trade-off in a meaningful way.

There are different ways to calculate the area under the P-R curve. Older methods used an 11-point interpolation, while modern methods (like the one used in the COCO challenge) compute the exact area using numerical integration.

4. Mean Average Precision (mAP): The Final Score

We now have the Average Precision (AP) for a single class. Mean Average Precision (mAP) is simply the mean of the APs calculated for all classes.

where is the number of classes.

But modern evaluation takes it one step further. Remember that our initial TP/FP classification depended on an IoU threshold (e.g., 0.5). A model that is "correct" at IoU=0.5 might be very imprecise at IoU=0.9. To reward models that are accurate at multiple levels of localization, the standard COCO metric calculates mAP over a range of IoU thresholds.

You will often see this written as mAP@[.5:.05:.95]. This means:

  1. Calculate the mAP with an IoU threshold of 0.5.
  2. Calculate the mAP with an IoU threshold of 0.55.
  3. ...and so on, up to an IoU threshold of 0.95.
  4. Finally, average all of these mAP scores.

This single number provides a very robust assessment of a model's overall performance.

Mean Average Precision (mAP) Explained and PyTorch Implementation

Let's return to Aladdin Persson's video to see how AP is extended to the full mAP metric.

Watch from 07:17 to 08:20. This section explains how AP scores are averaged across classes and, crucially, across different IoU thresholds to arrive at the final mAP score.

5. Implementation Deep Dive

Given your background, understanding how this is implemented in code will solidify the concepts. The algorithmic nature of mAP calculation—sorting, iterating, and tracking state—should feel familiar. The Aladdin Persson video continues with a full PyTorch implementation from scratch.

This is the most practical part of the lesson. We will focus on the key logic that translates the theory into code.

Mean Average Precision (mAP) Explained and PyTorch Implementation

Now, let's walk through a from-scratch PyTorch implementation of the mAP calculation. This will connect all the theoretical steps we've just discussed to concrete code.

Watch the implementation walkthrough from 08:20 to 26:01. You don't need to memorize every line, but focus on understanding the logic of these key parts: (11:00 - 15:00): The amount_bounding_boxes dictionary. This is a clever way to keep track of which ground truth boxes in each image have already been 'matched' with a high-confidence prediction. This ensures that only the best prediction for a given object is counted as a TP, while duplicates are correctly marked as FPs. (15:00 - 19:45): The main loop over detections. See how for each prediction, it finds the best matching ground truth box in the same image and checks if the IoU is above the threshold and if that ground truth box is still 'available'. (19:45 - 24:00): The use of torch.cumsum() on the true positive and false positive tensors. This is an efficient way to calculate the running TP and FP counts needed for the precision and recall values at each step of the sorted list. (24:00 - 25:00): The use of torch.trapz(). This function performs numerical integration using the trapezoidal rule to calculate the area under the precision-recall curve, giving us the final AP for the class.

Conclusion

You've now covered the entire pipeline for evaluating object detection and instance segmentation models, from the basic concept of IoU to the comprehensive mAP metric.

Key Takeaways:

  • IoU is the foundation, measuring the overlap between a single prediction and a ground truth box/mask.
  • An IoU threshold is used to classify detections as True Positives or False Positives.
  • Precision and Recall capture the trade-off between the quality and quantity of detections.
  • The Precision-Recall Curve visualizes this trade-off across all confidence scores.
  • Average Precision (AP) is the area under the P-R curve, providing a single-number summary for one class.
  • Mean Average Precision (mAP) is the final, robust metric, averaging AP across all classes and typically across a range of IoU thresholds (e.g., mAP@[.5:.05:.95]).

Preview of the next lesson:
We have now completed our deep dive into discriminative models for computer vision, which are trained to recognize patterns and make predictions. We are about to pivot to a completely different and exciting area of AI: generative models. Instead of just analyzing existing data, these models learn to create entirely new data. In our next lesson, we will begin this journey by implementing a Variational Autoencoder (VAE), one of the foundational generative architectures.

Can't find a good explanation? Sign up and we'll make it for you

Sign up