Introduction
Welcome to the first lesson of our second module, "Foundations: Serving and Profiling a Small Model." In the previous module, we built a complete mental model of the LLM inference process, culminating in understanding how to measure and diagnose performance using Time-to-First-Token (TTFT) and Inter-Token Latency (ITL).
Now, it's time to move from theory to practice. To benchmark, profile, and eventually optimize a model, you first need a working environment. This lesson will guide you through that foundational step. Our learning outcome is to set up a local inference environment with PyTorch and CUDA on a single consumer GPU.
We'll cover the essential software components, walk through a recommended installation process using Conda for robust environment management, and conclude by verifying that your GPU is correctly configured and accessible to PyTorch.
The GPU Software Stack: Driver, CUDA, and PyTorch
For PyTorch to leverage your NVIDIA GPU, a specific software stack must be in place. Understanding the role of each layer is crucial for both setup and troubleshooting.
- NVIDIA Driver: This is the lowest-level software that allows your operating system to communicate with the GPU hardware.
- CUDA Toolkit: This is NVIDIA's parallel computing platform and programming model. It provides the libraries and APIs necessary for general-purpose computing on GPUs.
- PyTorch: The deep learning framework itself. GPU-enabled PyTorch builds are compiled against specific versions of the CUDA Toolkit.
A common point of confusion is whether you need to manually install the full CUDA Toolkit from NVIDIA's website. For most use cases, the answer is no.
To clarify this critical point, let's consult a well-written guide on the topic.
Install PyTorch on Windows, Linux and macOS in 2025: Step by Step
The article 'Install PyTorch on Windows, Linux and macOS...' by Daya Shankar provides an excellent explanation of the relationship between the NVIDIA driver, the CUDA Toolkit, and PyTorch.
Please read the sections 'Unlocking GPU Power with NVIDIA CUDA' and the FAQ entry 'Do I Need the Full CUDA Toolkit from NVIDIA?'. Pay close attention to the key takeaway: you only need a compatible NVIDIA driver on your system. The PyTorch package itself will bundle the necessary CUDA runtime libraries.
The key insight from the reading is this: Your main responsibility is to ensure you have a sufficiently modern NVIDIA driver installed. The PyTorch installation process, when done correctly, handles the rest.
Step 1: Check Your NVIDIA Driver
Before installing anything else, open your terminal and run the NVIDIA System Management Interface (nvidia-smi) command:
nvidia-smi
If this command succeeds, you will see a table detailing your GPU, its power usage, memory, and, most importantly, the Driver Version and the CUDA Version. The CUDA version shown by nvidia-smi indicates the maximum version supported by your driver. You can install a PyTorch build compiled for this version or any earlier version.
If the command fails, you must install or update your NVIDIA drivers before proceeding.
Practical Guide: Setting Up Your Environment with Conda
While you can install these components manually or use Docker, we recommend using the Conda package manager for this course's development work. Conda excels at creating isolated environments and managing complex binary dependencies like the CUDA toolkit, which prevents conflicts and makes your setup more reliable.
The following video provides a complete walkthrough of creating a Conda environment and installing a GPU-enabled version of PyTorch.
Creating a conda Environment for Pytorch and Cuda
This video by Alex Soupir, 'Creating a conda Environment for Pytorch and Cuda,' clearly demonstrates the exact steps we will take to build our environment.
Watch the video from 01:30 to 08:26. It will guide you through: Creating a named Conda environment with a specific Python version. Installing the pytorch-cuda package, which lets Conda manage the CUDA toolkit dependency. Installing PyTorch, torchvision, and torchaudio from the correct channels. Verifying the installation in Python using torch.cuda.is_available().
Step-by-Step Instructions
Here is a summary of the process shown in the video, adapted with current best practices.
1. Create and Activate a Conda Environment:
It's essential to specify a Python version. Let's use Python 3.10, which is well-supported.
# Create a new environment named 'llm-systems' with Python 3.10
conda create --name llm-systems python=3.10
# Activate the environment
conda activate llm-systems
2. Install PyTorch with CUDA Support:
The best way to get the correct installation command is from the official PyTorch website.
- Go to the PyTorch Get Started page.
- Use the interactive tool to select your configuration:
- PyTorch Build: Stable
- Your OS: Linux (recommended), Windows, or Mac
- Package: Conda
- Language: Python
- Compute Platform: Select the CUDA version that is less than or equal to the one reported by
nvidia-smi. For example, ifnvidia-smishows CUDA 12.4, you can safely choose CUDA 12.1.
The website will generate a command. It will look something like this:
# Example command - DO NOT copy this blindly. Get the latest from the PyTorch website.
conda install pytorch torchvision torchaudio pytorch-cuda=12.1 -c pytorch -c nvidia
Run the command generated by the website in your activated llm-systems environment. This command instructs Conda to install PyTorch and its companion libraries, along with the specific CUDA toolkit version required, all neatly contained within your environment.
3. Verify the Installation:
This is the most critical step. You must confirm that PyTorch can see and use your GPU. Start a Python interpreter and run the following commands:
import torch
# 1. Check if CUDA is available
print(f"CUDA available: {torch.cuda.is_available()}")
# 2. If it is, print the CUDA version PyTorch was built with
if torch.cuda.is_available():
print(f"PyTorch CUDA version: {torch.version.cuda}")
# 3. Get the number of GPUs
gpu_count = torch.cuda.device_count()
print(f"Number of GPUs: {gpu_count}")
# 4. Get the name of the primary GPU
if gpu_count > 0:
print(f"GPU Name: {torch.cuda.get_device_name(0)}")
If torch.cuda.is_available() prints True and you see your GPU's name, your environment is correctly configured.
Platform-Specific Considerations
-
Linux (Recommended): This is the most mature and performant platform for deep learning. The steps above work seamlessly on distributions like Ubuntu.
-
Windows: For serious work on Windows, the Windows Subsystem for Linux (WSL2) is the standard. It allows you to run a full Linux environment directly on Windows, with deep integration for GPU access. The manual installation process for native Windows can be fragile.
)
This diagram from NVIDIA shows the architecture of GPU access in WSL2. The Windows NVIDIA driver is shared with a Linux kernel, allowing you to run a native Linux environment with full CUDA capabilities. -
macOS: Newer Apple Silicon Macs use the Metal Performance Shaders (MPS) backend, not CUDA. The setup is
conda install pytorch torchvision torchaudio -c pytorchand verification istorch.backends.mps.is_available(). While functional, this course will focus on the CUDA ecosystem, as it dominates production GPU serving.
Troubleshooting
If torch.cuda.is_available() returns False, it's a classic setup issue. The troubleshooting table in the Install PyTorch... article you read earlier is an excellent resource. Here are the most common culprits:
- Environment Mismatch: You are running Python from an environment other than the one where you installed the GPU-enabled PyTorch. Ensure
conda activate llm-systemsis active. - Incorrect PyTorch Package: You may have accidentally installed the CPU-only version of PyTorch. Re-run the installation command from the PyTorch website, making sure to select a CUDA compute platform.
- NVIDIA Driver Issue: Your driver might be too old for the CUDA version you are trying to use. Check
nvidia-smiand, if necessary, update your drivers from the NVIDIA website.
Conclusion
You have now built the foundational layer of our hands-on work: a dedicated, GPU-accelerated Python environment. This clean, isolated workspace is where we will build, profile, and optimize our models throughout the course.
Key Takeaways:
- The GPU software stack consists of the NVIDIA Driver, CUDA libraries, and PyTorch.
- For most modern workflows, you only need to manage the NVIDIA driver at the system level. The required CUDA libraries are bundled with the PyTorch package.
- Conda is a powerful tool for creating isolated environments and managing these complex dependencies, preventing conflicts.
- The ultimate test of your setup is a successful call to
torch.cuda.is_available()within your activated Conda environment.
Preview of the Next Lesson:
With our environment ready, we can finally get a model up and running. In the next lesson, we will load and serve a small model (e.g., Phi-3-mini) using Hugging Face Transformers. This will be our first practical step in observing a real LLM in action on our local machine.