Introduction
Welcome to the final lesson in our module on production inference frameworks. In the previous lesson, you built a decision matrix to strategically select between vLLM, SGLang, and TensorRT-LLM based on specific workload characteristics. You've moved from being a user of these frameworks to an architect who can make informed system design choices.
Today, we put that architectural thinking into practice with a full-scale deployment exercise. Your learning outcome is to deploy a 70B+ model on a multi-GPU setup using a chosen framework and load-test it with 500 concurrent users. This lesson synthesizes everything you've learned about inference engines, hardware, and performance metrics into a realistic, end-to-end production scenario.
We will tackle this by:
- Choosing a framework and model appropriate for a high-concurrency workload.
- Walking through the steps to deploy this model across multiple GPUs using tensor parallelism.
- Defining a methodology and using industry-standard tools to load-test the deployment against our target of 500 concurrent users.
- Analyzing the results to understand the system's performance and identify potential tuning opportunities.
This is the capstone exercise for this module and directly prepares you for the final project, where you will benchmark your own custom-built engine against these production-grade systems.
1. The Scenario: Framework and Model Selection
Let's define our production goal: We need to build an API for a Llama 3 70B model that can handle up to 500 concurrent users. The application is an interactive chat service, so while overall throughput is important for cost-efficiency, user-perceived latency (especially Time-To-First-Token) must be low.
Framework Choice:
Referring back to the decision matrix from our last lesson, we need a framework that excels at high-concurrency throughput.
- TensorRT-LLM is a strong contender for raw performance but its high operational cost (compilation, C++ focus) makes it less ideal for rapid deployment and iteration.
- SGLang offers excellent performance and flexibility, especially for structured output and workloads with high prefix sharing. Its RadixAttention and cache-aware scheduler make it a top-tier choice.
- vLLM is renowned as the "high-concurrency champion," designed specifically to maximize throughput with its PagedAttention mechanism. It's Python-native and incredibly easy to get started with.
Given our primary constraint is high concurrency and ease of deployment, both vLLM and SGLang are excellent choices. For this lesson, we will proceed with vLLM due to its widespread adoption and clear documentation for this exact use case. However, the principles and testing methodologies we discuss are directly transferable to an SGLang deployment.
Model Choice:
We will use nvidia/Llama-3.3-70B-Instruct-FP8. The 70B parameter size makes it a powerful model, while the FP8 quantization significantly reduces the VRAM footprint from ~140GB (in FP16) to ~70GB, making it feasible to serve on a multi-GPU setup with common cloud instances (e.g., 2x A100 40GB or 4x L40S 48GB).
2. Deploying a 70B Model with vLLM
Deploying a model that exceeds the memory of a single GPU requires parallelism. The most common strategy for this is Tensor Parallelism (TP), where the weight matrices of the model are sharded (split) across multiple GPUs. vLLM makes this straightforward.
The following guide is an excellent, practical walkthrough for deploying this exact model. We will follow its key steps.
Quick Start Recipe for Llama 3.3 70B on vLLM
This document from the vLLM project provides a complete recipe for deploying the Llama 3.3 70B model on NVIDIA GPUs. It covers everything from Docker setup to multi-GPU server launch commands and performance tuning.
Read through the 'Introduction', 'Deployment Steps', and 'Configs and Parameters' sections. Focus on understanding the sequence of actions: pulling the Docker image, launching the container, and, most importantly, the vllm serve command with the --tensor-parallel-size flag. Pay close attention to the explanation of tunable parameters like max-num-seqs and max-num-batched-tokens, as these are the levers you'll pull to optimize performance.
Key Deployment Steps Summary
Let's assume you are on a machine with at least two NVIDIA Hopper or Ampere GPUs and have the NVIDIA Container Toolkit installed. Based on the guide, the process is as follows:
-
Pull the Docker Image: Get the pre-built vLLM container, which includes all necessary dependencies.
docker pull vllm/vllm-openai:latest -
Run the Container: Start a container with access to all available GPUs.
docker run --gpus all --ipc=host --rm -it vllm/vllm-openai:latest /bin/bash--gpus all: Makes the GPUs available inside the container.--ipc=host: Allows processes (like the multiple GPU workers) to communicate efficiently via shared memory.
-
Launch the vLLM Server: This is the critical command. Inside the container, you launch the server, telling it to use tensor parallelism to split the model across GPUs. To run a 70B FP8 model, you need at least 70GB of VRAM. If you're on a machine with 4x L40S (48GB each), you can use a tensor parallelism size of 2, 3 or 4. Let's use
TP=2to start, which would place ~35GB on each of two GPUs.vllm serve nvidia/Llama-3.3-70B-Instruct-FP8 \ --tensor-parallel-size 2 \ --kv-cache-dtype fp8 \ --host 0.0.0.0 \ --port 8000--tensor-parallel-size 2: This is the key. It instructs vLLM to partition the model's weights across 2 GPUs.--kv-cache-dtype fp8: Using FP8 for the KV cache further reduces memory pressure, allowing for larger batches.--host 0.0.0.0: Exposes the server outside the Docker container.
After a few minutes of downloading and loading the model shards onto the GPUs, you'll see a message indicating the server is ready to accept requests at http://<your-ip>:8000. You now have a 70B model served across multiple GPUs.
3. Load Testing at Scale
With the server running, how do we verify it can handle 500 concurrent users? Sending requests in a simple for loop is insufficient as it doesn't simulate true concurrency. We need a dedicated load testing tool.
While tools like wrk or ApacheBench are great for standard web servers, LLM APIs are stateful and have long-running streaming responses, requiring specialized tools.
Option 1: vLLM's Built-in Benchmark Tool
For a quick and easy test, vLLM provides a benchmarking script. It's a great first step to get a feel for your server's limits.
Quick Start Recipe for Llama 3.3 70B on vLLM
The same vLLM recipe we used for deployment also details how to use its built-in benchmarking tool. This is the simplest way to apply a concurrent load.
Read the 'Benchmarking Performance' and 'Interpreting Performance Benchmarking Output' sections. Note the --max-concurrency flag, which is how you'll simulate your 500 users.
You would run a command like this from a separate terminal:
vllm bench serve \
--host <your-ip> \
--port 8000 \
--model nvidia/Llama-3.3-70B-Instruct-FP8 \
--dataset-name random \
--random-input-len 1024 \
--random-output-len 1024 \
--max-concurrency 512 \
--num-prompts 2560
This script will fire requests to your server, maintaining up to 512 in-flight requests, and then report metrics like throughput (tok/s), TTFT, and TPOT.
Option 2: A Professional-Grade Tool (AI Perf)
For more sophisticated analysis, especially around meeting specific Service Level Objectives (SLOs), a tool like NVIDIA's AI Perf is invaluable. It introduces concepts essential for production monitoring.
AI Perf benchmarking - Dynamo and other LLM endpoints
This video from NVIDIA Developer introduces AI Perf, a tool designed for benchmarking large-scale LLM deployments. It covers how to simulate concurrency and, most importantly, how to measure 'goodput'—the percentage of requests that meet your performance targets.
Watch the sections on running a simple benchmark (05:39 - 12:46), goodput analysis (30:32 - 33:22), and prioritizing metrics (38:27 - 41:17). Focus on how you define concurrency (--concurrency) and how goodput analysis (--goodput) transforms raw numbers into a business-relevant metric.
Using a tool like AI Perf, you can simulate the 500-user load and evaluate performance against an SLO. For example: "99% of requests must have a TTFT under 500ms."
A hypothetical AI Perf command would look like this:
aiprf profile \
--url http://<your-ip>:8000/v1/chat/completions \
--model nvidia/Llama-3.3-70B-Instruct-FP8 \
--concurrency 500 \
--request-count 5000 \
--input-len 1024 \
--output-len 1024 \
--goodput "ttft_ms:500"
The output wouldn't just give you average latencies; it would tell you that your Total Throughput was X req/s, but your Goodput was Y req/s, meaning only a fraction of requests met the 500ms TTFT target. This is a much more powerful insight than averages alone.
4. Analyzing Results and Tuning the System
Running a load test is the beginning, not the end. The real work of an AI Systems Engineer is interpreting the results and tuning the system.
Let's watch a segment of this video which provides realistic performance numbers for different multi-GPU configurations and discusses the trade-offs.
How to pick a GPU and Inference Engine?
This video from Trelis Research benchmarks various Llama models on different multi-GPU setups using SGLang. While the framework is different, the performance characteristics and cost-benefit analysis are highly relevant for understanding what to expect from our vLLM deployment.
Watch the section on comparing costs and performance (43:13 - 01:00:16). Pay attention to how performance (tokens/sec) changes when moving from 2 to 4 GPUs for a 70B model. Also, note the discussion on balancing cost vs. customer experience (latency), which is the central challenge in production serving.
Interpreting Hypothetical Results:
Imagine our load test with 500 concurrent users yields the following:
- Output Token Throughput: 15,000 tok/s
- Median TTFT: 450 ms
- P99 TTFT: 3,500 ms
- Error Rate: 2%
Analysis:
The system is handling a massive load (15k tok/s is very high!), and the typical user has a good experience (450ms TTFT). However, the P99 TTFT is terrible (3.5s), and we're seeing some errors. This suggests the server is becoming overloaded, and requests are getting stuck in a queue, leading to timeouts for some users. The server is configured for maximum throughput, but at the expense of latency spikes under pressure.
Tuning Levers:
How do we fix this? We can adjust vLLM's serving parameters.
Quick Start Recipe for Llama 3.3 70B on vLLM
Let's revisit the vLLM recipe, this time focusing on the parameters that allow us to balance throughput and latency.
Read the final section, 'Balancing between Throughput and Latencies'. This is the most important part for system tuning. Understand the trade-offs of adjusting --tensor-parallel-size and --max-num-seqs.
Based on this, our tuning strategy could be:
- Reduce
max-num-seqs: This parameter controls the maximum number of sequences (requests) in a batch. Our high P99 latency suggests the current value might be too high, allowing too many requests to be batched and causing some to wait too long. We could relaunch the server with--max-num-seqs 256(down from a high default) to enforce smaller, faster-moving batches. This will likely reduce total throughput but should improve the P99 latency. - Increase
tensor-parallel-size: The guide notes that higher TP typically results in lower latencies. If we have 4 GPUs, we could try running withTP=4instead ofTP=2. This increases communication overhead and might lower per-GPU throughput, but it could significantly improve our P99 TTFT by reducing the computation time for each step.
After making a change, we would run the load test again, compare the metrics, and iterate until we find a configuration that meets our SLOs for both throughput and latency.
Conclusion
In this lesson, you have walked through the complete lifecycle of deploying and stress-testing a large language model in a production-like setting. This is the practical application of all the theoretical knowledge you've accumulated.
Key Takeaways:
- Large Model Deployment Requires Parallelism: You cannot serve a 70B+ model on a single consumer GPU. Tensor Parallelism is the standard technique, and frameworks like vLLM and SGLang make it accessible via simple command-line flags.
- Load Testing is Non-Negotiable: You cannot claim a system is "performant" or "scalable" without measuring it under a realistic, concurrent load. Tools like
vllm bench serveorAI Perfare essential for this. - Averages are Deceiving; Percentiles and Goodput Tell the Story: Focusing on P99 latency and goodput (requests meeting an SLA) is critical for understanding the actual user experience under load.
- Tuning is an Iterative Process: Performance engineering is a cycle of deploying, measuring, analyzing, and tuning. The key is knowing which parameters (
max-num-seqs,tensor-parallel-size, etc.) to adjust to balance the fundamental trade-off between throughput and latency.
Preview of the Next Module:
You have now reached the end of the modules focused on using and evaluating existing frameworks. In the upcoming capstone module, you will take the ultimate step in becoming an AI Systems Engineer: you will build your own custom inference engine from scratch. You will implement many of the concepts we've discussed—continuous batching, paged KV caching, and more—and then use the benchmarking skills you learned today to measure your engine's performance against the very tools you've just mastered.