Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 9
Distributed Inference for Large Models
1
PyTorch Distributed: NCCL Backend for Multi-GPU Communication
Set up a PyTorch distributed process group using the NCCL backend for multi-GPU communication.
2
Tensor Parallelism: Sharding Weights for Parallel Layer Execution
Explain tensor parallelism and how it shards weight matrices across GPUs to run a single layer in parallel.
3
Tensor-Parallel Transformer Inference
Implement tensor-parallel inference for a transformer block across 2+ GPUs, including the necessary all-reduce communication steps.
4
Pipeline Parallelism: Distributing Model Layers Across GPUs
Explain pipeline parallelism and how it partitions model layers across multiple GPUs.
5
Implementing Forward Pass in Pipeline-Parallel Models
Implement a forward pass for a pipeline-parallel model, managing micro-batches and data transfers between stages.
6
Visualizing Pipeline Bubbles in GPU Utilization
Visualize the GPU utilization timeline for a pipeline-parallel implementation to identify and measure the 'pipeline bubble'.
7
Communication Overhead in Parallelism
Analyze the communication overhead (all-reduce vs. point-to-point) of tensor vs. pipeline parallelism.
8
Optimizing 70B Model Serving on 4xH100: Memory & Communication Trade-offs
Determine an optimal parallelism configuration (e.g., hybrid TP/PP) for serving a 70B model on a 4xH100 cluster based on memory and communication trade-offs.
Previous module
High-Throughput Serving: Batching and Scheduling
Next module
Evaluating Production Inference Frameworks