Introduction
In our last lesson, we saw how torch.compile can dramatically improve performance by fusing operations and generating optimized Triton kernels. We used torch.profiler to verify this, observing a shift from many small, distinct aten kernels to a few large, fused ones. This confirmed that an optimization occurred.
However, torch.profiler and its Chrome trace view stop at the kernel boundary. They tell you which kernels ran and for how long, but they don't tell you how efficiently those kernels used the GPU's resources. Was the fused kernel compute-bound or memory-bound? Was it stalled waiting for data? Did it suffer from inefficient memory access patterns?
To answer these questions, we need to move beyond PyTorch-level profilers and use tools that operate at the CUDA driver level. This lesson introduces NVIDIA's Nsight suite, the professional standard for GPU profiling. Your objective is to learn how to use Nsight Systems to get a system-wide overview and then use Nsight Compute to perform a deep-dive analysis of individual kernels, allowing you to pinpoint the precise reasons for performance bottlenecks.
The Profiler's Toolkit: Nsight Systems vs. Nsight Compute
NVIDIA provides two primary profiling tools, and it's crucial to understand their distinct roles. They are designed to be used together in a top-down analysis workflow.
Reimplementing FlashAttention for performance and giggles
This article, while focused on reimplementing FlashAttention, provides an excellent, concise breakdown of the different profiling tools and their specific use cases.
Read just the 'Profiling Tools' section. The author clearly distinguishes between torch.profiler, Nsight Systems (nsys), and Nsight Compute (ncu), which perfectly sets the stage for this lesson.
To summarize the distinction:
-
NVIDIA Nsight Systems (
nsys): This is a system-level profiler. It captures a timeline of all major events on both the CPU and GPU: CUDA API calls, kernel launches, and memory transfers between host (CPU) and device (GPU). Its main purpose is to identify system-level bottlenecks.- Use it to answer: "Where is time being spent overall?", "Is my GPU idle because the CPU can't feed it fast enough?", "Which 2-3 kernels are consuming the most GPU time?"
-
NVIDIA Nsight Compute (
ncu): This is a kernel-level profiler. It focuses on a single kernel (or even a specific invocation of it) and provides an exhaustive analysis of its execution, including detailed performance counters from the hardware itself.- Use it to answer: "Why is this specific kernel slow?", "Is this kernel limited by compute or memory bandwidth?", "What are the warp stall reasons?", "Are there shared memory bank conflicts?"
The standard workflow is to first use nsys to find the most expensive kernels and then use ncu to analyze them in detail.
Step 1: System-Level Analysis with Nsight Systems (nsys)
Let's start with nsys. You can run it from the command line, wrapping the Python script you want to profile.
A typical command looks like this:nsys profile -o <report_name> python my_script.py
The output is a .nsys-rep file that can be viewed in the Nsight Systems GUI or analyzed on the command line.
A Practical Example: Profiling vLLM
For a realistic use case, let's see how you would profile a state-of-the-art inference server like vLLM.
The vLLM documentation provides a practical guide for profiling their system, which includes recommended flags and commands for nsys.
Please review this document, paying attention to three key parts: 'Profile with NVIDIA Nsight Systems': Note the installation and the recommended flags like --trace-fork-before-exec=true. 'Offline Inference': Look at the example command. You can see how nsys profile is simply prefixed to a standard vLLM benchmark command. This is the basic usage pattern. 'Analysis': Examine the CLI example output under 'CUDA GPU Kernel Summary'. This table, sorted by Time (%), is the primary output you'll use from nsys. It directly tells you which kernels are your top offenders.
The key takeaway from the vLLM guide is that nsys immediately gives you a ranked list of the most time-consuming kernels. This is your hit list for optimization.
Visualizing the Timeline
While the CLI summary is useful, the nsys GUI provides a rich visual timeline that helps you understand the dynamics of the system.

When you analyze a timeline like this, you're looking for two things:
- Long Bars: Which kernels are taking the most time? This visually corresponds to the CLI summary table.
- Gaps: Are there periods where the GPU stream is empty? This indicates that the GPU is idle, likely waiting for the CPU to launch the next kernel. This was the problem
torch.compilehelped solve in the last lesson.
For a hands-on demonstration of running nsys and interpreting the visual trace, watch the following clip.
Lecture 16: On Hands Profiling
The following video shows a presenter running nsys on a model and analyzing the output.
Watch from 04:26 to 06:42. The presenter runs nsys and opens the trace. Notice how he immediately identifies the startup phase versus the steady state, and then zooms in on the CUDA kernels to see what's running. He points out that while you can see kernel names, it's hard to get prescriptive details from this view alone, which perfectly motivates the need for ncu.
Once nsys has helped you identify your target kernel (e.g., sm90_xmma_gemm... from the image), it's time to bring out the microscope: Nsight Compute.
Step 2: Deep-Dive Kernel Analysis with Nsight Compute (ncu)
ncu provides an unparalleled view into the inner workings of a single kernel's execution. It tells you not just how long it took, but how it used the hardware's resources.
A basic ncu command specifies the target kernel to profile:ncu --kernel-name <regex_for_kernel_name> --set full -o <report_name> -f python my_script.py
--kernel-name: Focuses the analysis on specific kernels. This is crucial as profiling all kernels is extremely slow.--set full: Collects all available metrics.-f: Forces the output file to be overwritten.-o: Specifies the output.ncu-repfile.
The Roofline Model: Your First Check
One of the most powerful visualizations in ncu is the roofline model. It immediately tells you whether your kernel's performance is limited by the GPU's computational power or its memory bandwidth.

- Arithmetic Intensity: The ratio of floating-point operations to bytes of data moved.
- Memory-Bound: Performance is limited by how fast you can feed the compute units with data from memory. To improve, you need to reduce data movement.
- Compute-Bound: Performance is limited by the raw processing power of the SMs. To improve, you need to use more efficient instructions or increase parallelism.
Case Study: Diagnosing Bottlenecks in FlashAttention
To see ncu in action, we'll follow a real-world investigation from an expert developer who reimplemented FlashAttention and used ncu to systematically find and fix bottlenecks. This is a masterclass in kernel profiling.
Reimplementing FlashAttention for performance and giggles
We will use the article 'Reimplementing FlashAttention for performance and giggles' as our case study. It's a fantastic, deep, and practical walkthrough of the profiling workflow.
This is the core of the lesson. Please read these sections carefully. We will walk through the author's optimization journey step-by-step: 'Profiling the v1 Implementation' (Section 1): The author profiles their first attempt. ncu immediately points out two problems: massive HBM traffic (confirming it's memory-bound) and low occupancy due to high shared memory usage. This is the first layer of analysis. 'Profiling the v2 Implementation' (Section 2): After rewriting the kernel to reduce HBM traffic, the speedup is disappointing. ncu reveals the new bottleneck: Shared Memory Bank Conflicts and Uncoalesced Accesses. The problem has shifted from global memory to on-chip memory. 'Understanding the Shared Memory Bottleneck' (Section 3): This is a crucial theoretical detour. The author explains what bank conflicts and wavefronts are, directly interpreting the metrics reported by ncu. This connects the high-level symptom (slowness) to a specific hardware mechanism. 'PTX and fun' (Section 4): To confirm the hypothesis, the author dives into the generated PTX (assembly) code to find the exact strided memory access pattern causing the bank conflicts. This is the deepest level of analysis, linking the Python code to the hardware's behavior.
This case study perfectly illustrates the iterative process of profiling:
- Profile with
ncu. - Identify the primary bottleneck (e.g., HBM bandwidth).
- Refactor the code to address it.
- Profile again.
- Identify the next bottleneck (e.g., shared memory bank conflicts).
- Repeat.
This process continues until performance is satisfactory or you are verifiably hitting the theoretical limits of the hardware. ncu is the tool that makes this data-driven approach possible.
Another common bottleneck you'll encounter are host-device synchronizations, which can destroy performance by forcing the GPU to wait. The following video explains how to spot and eliminate them.
Lecture 16: On Hands Profiling
Even with optimized kernels, simple mistakes in your Python code can create performance cliffs. This video explains how to find cudaStreamSynchronize calls, which are a common culprit.
Watch from 42:14 to 48:58. The presenter uses the PyTorch profiler's trace to find an unexpected cudaStreamSynchronize call. Notice how a seemingly innocuous Python operation (max(1, x)) forces a synchronization because it converts a tensor to a Python scalar. This forces the CPU to wait for the GPU result, creating a major stall. Nsight Systems would show this as a large gap in the GPU timeline. The fix is to keep operations in the tensor domain (torch.max(torch.ones_like(x), x)).
Conclusion
You have now moved from a high-level understanding of GPU execution to the most granular level of kernel analysis. You're equipped with the same tools and workflow used by performance engineers at NVIDIA and in top AI labs.
Key Takeaways:
- Top-Down Profiling: Use Nsight Systems (
nsys) first to get a system-wide view and identify the most time-consuming kernels. - Deep-Dive Analysis: Use Nsight Compute (
ncu) to analyze those top kernels. Start with the roofline model to determine if they are compute- or memory-bound. - Iterative Optimization:
ncuprovides specific, actionable metrics (e.g., memory traffic, occupancy, warp stalls, bank conflicts) that guide your code refactoring. The process is a loop: profile, analyze, fix, and repeat. - Connecting Code to Hardware: The ultimate goal is to understand how your high-level Python/Triton code translates into low-level hardware behavior. Tools like
ncuand the ability to inspect PTX assembly make this connection explicit.
Preview of the Next Lesson:
Our deep dive into profiling has repeatedly shown that for LLMs, performance is often dominated by memory access patterns. In the next module, we will focus on the most critical component of a transformer in this regard: the attention mechanism. We will begin by profiling the standard scaled dot-product attention and using ncu to identify its memory-access bottleneck. This will set the stage for understanding FlashAttention, an algorithm specifically designed to be "profiler-friendly" by minimizing costly reads and writes to HBM.