Skip to main content
Create your own

Non-Maximum Suppression for Object Detection

Hello! Welcome to the next lesson in our exploration of advanced computer vision.

Introduction

In our previous lesson, we dove into single-shot object detectors like YOLO and SSD. We saw how these models achieve real-time performance by producing a dense grid of predictions in a single forward pass. However, this process results in thousands of raw, overlapping bounding boxes for potentially just a few objects in an image. This is not the clean, final output we need.

Today, we will tackle this problem head-on. Your learning outcome is to apply non-maximum suppression (NMS) to refine detection bounding boxes. NMS is the essential post-processing algorithm that filters this noisy output, discarding redundant boxes and keeping only the most confident detection for each object.

We will cover:

  1. The core problem that NMS solves.
  2. The key metric it relies on: Intersection over Union (IoU).
  3. The step-by-step NMS algorithm.
  4. How to implement NMS, both from scratch and using standard libraries.
  5. The importance of its hyperparameters and potential pitfalls.

By the end of this lesson, you'll understand how raw model outputs are transformed into the final, clean detections you see in applications.

1. The Problem: A Flood of Detections

Object detection models like YOLO and SSD don't just output one box per object. They evaluate thousands of potential locations, scales, and aspect ratios (priors or anchor boxes), resulting in multiple, highly overlapping detections for a single object.

Non-Maximum Suppression (NMS) Demonstration
This image perfectly illustrates the problem and the solution. On the left, the raw output from a detector shows multiple overlapping boxes for the dog, bicycle, and truck. On the right, after applying NMS, we have a single, clean detection for each object.

The goal is to suppress the "non-maximal" boxes—the ones that are redundant and have lower confidence scores—and keep only the best one. Let's watch a brief video that introduces this concept.

Non Max Suppression Explained and PyTorch Implementation

The video 'Non Max Suppression Explained and PyTorch Implementation' by Aladdin Persson clearly sets up the problem that NMS is designed to solve.

Watch the first 1 minute and 18 seconds of the video. It provides a great high-level intuition for why we need a 'cleanup' method for our bounding box predictions.

To perform this cleanup, the algorithm needs a way to quantify how much two boxes overlap. This brings us to a fundamental metric in object detection.

2. The Core Metric: Intersection over Union (IoU)

Intersection over Union (IoU), also known as the Jaccard index, is a simple and powerful metric that measures the extent of overlap between two bounding boxes. It's calculated by dividing the area of the overlap (the intersection) by the area of the combined boxes (the union).

The formula is:

The IoU value ranges from 0 (no overlap) to 1 (perfect overlap).

Bounding Box Intersection Calculation for IoU
This diagram shows how to calculate the coordinates of the intersection box. Given two boxes, A and B, the top-left corner of the intersection is found by taking the maximum of their top-left coordinates, and the bottom-right corner is found by taking the minimum of their bottom-right coordinates.

To learn more about IoU and its calculation, please read the following section from an article by LearnOpenCV.

Intersection Over Union (IoU)

This section of the article 'Non Maximum Suppression: Theory and Implementation in PyTorch' provides a clear definition of IoU and includes a Python implementation that you'll find very familiar.

Read the section titled 'Intersection Over Union (IoU)'. Pay attention to the formula and the visual representation of intersection and union areas.

Now that we have a way to measure overlap, we can formally define the NMS algorithm.

3. The Non-Maximum Suppression Algorithm

NMS is a greedy, iterative algorithm that works on a per-class basis. Here's a breakdown of the steps:

  1. Select a Class: The entire process is performed independently for each object class predicted by the model (e.g., first for all "car" boxes, then for all "person" boxes, etc.).

  2. Filter by Confidence: Start with a list of all predicted boxes for the selected class. It's common practice to first discard any boxes with a confidence score below a certain score_threshold (e.g., 0.5). This is a preliminary filtering step.

  3. Sort by Score: Sort the remaining bounding boxes in descending order based on their confidence scores.

  4. Iterate and Suppress:
    a. Select the bounding box with the highest confidence score and move it to a final list of "kept" boxes.
    b. Calculate the IoU of this highest-scoring box with all other boxes remaining in the sorted list.
    c. Discard (suppress) any box from the list that has an IoU greater than a predefined iou_threshold.
    d. Return to step 4a, taking the next-highest-scoring box from the now-reduced list.

  5. Repeat: Continue this process until the list of boxes is empty. The "kept" list contains the final detections for that class.

This video walks through a numerical example of this process.

(NMS) Non Maximum Suppression explained in detail using example. NMS algorithm explained.

The video '(NMS) Non Maximum Suppression explained in detail using example' from Datum Learning provides a clear, step-by-step example of the algorithm in action.

Watch from 02:19 to 05:19. Follow along as the presenter applies the NMS algorithm to a list of candidate boxes, showing how boxes are selected and suppressed based on confidence and IoU.

Test your understanding!

Imagine you have the following bounding boxes for the class "dog":

  • Box A: score = 0.95
  • Box B: score = 0.90
  • Box C: score = 0.75
  • Box D: score = 0.70

The IoU values are:

  • IoU(A, B) = 0.85
  • IoU(A, C) = 0.10
  • IoU(B, C) = 0.05
  • IoU(A, D) = 0.80
  • IoU(C, D) = 0.08

If your iou_threshold is 0.5, which boxes will be kept after applying NMS?

Show answer
  1. The boxes are already sorted by score: [A, B, C, D].
  2. Select Box A (score 0.95). Add it to the keep list.
  3. Calculate IoU with other boxes:
    • IoU(A, B) = 0.85 > 0.5. Suppress Box B.
    • IoU(A, C) = 0.10 < 0.5. Keep Box C for now.
    • IoU(A, D) = 0.80 > 0.5. Suppress Box D.
  4. The remaining list of boxes to consider is just [C].
  5. Select Box C (score 0.75). Add it to the keep list.
  6. The list is now empty.

The final kept boxes are A and C. This scenario represents two distinct dogs that were correctly detected.

4. Implementation in Practice

Understanding the algorithm is one thing; implementing it efficiently is another. Given your background, you'll appreciate seeing how this logic translates into code.

From-Scratch Implementation in PyTorch

A from-scratch implementation is the best way to solidify your understanding. The logic involves sorting, a while loop, and careful tensor indexing to filter out suppressed boxes.

Non Max Suppression Explained and PyTorch Implementation

Let's return to Aladdin Persson's video. He provides a clean PyTorch implementation that directly translates the algorithm we just discussed into code.

Watch from 04:48 to 10:57. Focus on the core logic inside the while loop: Boxes are pre-sorted by confidence score. The highest-scoring box is popped from the list (chosen_box). A list comprehension is used to build a new boxes list, keeping only those that either belong to a different class or have an IoU with chosen_box that is less than the threshold. The chosen_box is appended to the final results.

This implementation is a great example of translating an algorithm into efficient, vectorized code, a common task in ML engineering.

Using Pre-built Library Functions

In most production code, you wouldn't write NMS from scratch. Optimized and battle-tested versions are available in major libraries like OpenCV and PyTorch/Torchvision.

The article Non-Maximum Suppression with OpenCV and Python shows how to use OpenCV's cv2.dnn.NMSBoxes() function. This is a very common approach.

The key parameters are:

  • bboxes: A list of bounding box coordinates.
  • scores: A list of corresponding confidence scores.
  • score_threshold: The confidence threshold for the initial filtering step. Any box with a score below this is discarded immediately.
  • nms_threshold: This is our iou_threshold. Any two boxes with an IoU above this are candidates for suppression.

The function returns the indices of the boxes that should be kept.

5. Hyperparameters and Pitfalls

NMS is a simple and effective algorithm, but it's not perfect. Its performance is highly dependent on the choice of the iou_threshold.

  • If iou_threshold is too low (e.g., 0.3): The algorithm becomes very aggressive. It might suppress correct detections of distinct, but heavily overlapping, objects. Imagine a photo of a dense crowd where people are partially occluding each other.
  • If iou_threshold is too high (e.g., 0.9): The algorithm becomes too permissive, and you may end up with multiple, highly overlapping boxes for the same object, defeating the purpose of NMS.

Let's watch a brief segment that highlights this exact problem.

(NMS) Non Maximum Suppression explained in detail using example. NMS algorithm explained.

The Datum Learning video points out a classic failure case of NMS when objects are close together.

Watch from 05:19 to 06:40. The video shows an example where two different objects (a person and a horse) are very close. A standard IoU threshold might incorrectly suppress one of them.

This highlights that the iou_threshold is a crucial hyperparameter that often needs to be tuned for a specific dataset or application. More advanced techniques like Soft-NMS have been developed to address this by decaying the scores of overlapping boxes instead of outright eliminating them, but standard NMS remains the most widely used method due to its simplicity and speed.

Conclusion

In this lesson, we demystified the crucial post-processing step that turns a chaotic mess of predictions into a clean set of object detections.

Key Takeaways:

  • Purpose: Non-Maximum Suppression (NMS) is a post-processing algorithm that filters the dense, overlapping bounding boxes from an object detector, keeping only the most confident prediction for each object.
  • Core Metric: NMS relies on Intersection over Union (IoU) to measure the degree of overlap between boxes.
  • Algorithm: It is a greedy, iterative process that, for each class, repeatedly selects the highest-scoring box and suppresses all other boxes that overlap with it above a certain iou_threshold.
  • Implementation: You can implement NMS from scratch for educational purposes, but in practice, optimized library functions (like cv2.dnn.NMSBoxes or torchvision.ops.nms) are preferred.
  • Hyperparameter: The iou_threshold is a critical hyperparameter that balances the trade-off between suppressing duplicates and incorrectly removing valid detections of distinct, overlapping objects.

Preview of the next lesson:
We now have a complete pipeline: a model predicts thousands of boxes, and NMS cleans them up to produce a final set of detections. But how do we measure how good these final detections are? In our next lesson, we will focus on evaluating object detection models using appropriate metrics, such as the mean Average Precision (mAP), which itself relies on the concept of IoU.

Can't find a good explanation? Sign up and we'll make it for you

Sign up