Skip to main content
Create your own

Benchmarking Custom Engine vs. vLLM: Performance and Architectural Analysis

Introduction

Welcome to the final lesson of the capstone project. In the previous lesson, you successfully implemented the request ingestion and response streaming layers, transforming your collection of backend components into a fully functional, end-to-end LLM inference service. Your custom engine is now architecturally complete, capable of accepting and processing requests just like a production system.

This brings us to the ultimate test. The learning outcome for this lesson is to benchmark the custom engine against vLLM under a realistic traffic pattern and analyze its architectural strengths and weaknesses. You will not just measure performance, but dissect the results to understand why your engine behaves the way it does compared to a state-of-the-art framework. This process of benchmarking, analysis, and architectural reasoning is a cornerstone of AI Systems Engineering.

We will cover four key areas:

  1. Defining the Benchmark: What key metrics matter for LLM inference?
  2. Designing the Workload: How do we create a "realistic traffic pattern"?
  3. Running the Test: Using a professional benchmarking tool to test your engine and vLLM.
  4. Analyzing the Results: Connecting performance data back to architectural decisions to identify the strengths and weaknesses of your design.

1. Defining the Benchmark: Metrics and Traffic Patterns

A benchmark is only as meaningful as the metrics you collect and the workload you simulate. Let's establish a clear framework for both.

Key Performance Metrics

When evaluating an LLM serving system, we are interested in more than just raw speed. We need to measure different aspects of the user experience and system capacity.

Mastering LLM Inference Optimization From Theory to Cost Effective Deployment: Mark Moyou

To start, let's get a concise overview of the three most important metrics for LLM inference. This video by Mark Moyou provides a clear explanation of each.

Watch the segment from 15:09 to 16:25. Focus on the definitions of: Time to First Token (TTFT): Measures the performance of the prefill stage. Inter-Token Latency (ITL): Measures the performance of the autoregressive decode steps. Throughput: Often measured in output tokens per second, indicating overall system capacity.

These metrics allow us to diagnose performance bottlenecks. High TTFT points to an inefficient prefill phase, while high ITL suggests a slow decoding loop. Throughput tells us how the system performs under concurrent load.

Designing a Realistic Traffic Pattern

Running a benchmark with a single, uniform type of request doesn't reflect real-world usage. A production server handles a mix of queries with varying input and output lengths.

Mastering LLM Inference Optimization From Theory to Cost Effective Deployment: Mark Moyou

Let's continue with the same video to understand the different query patterns that comprise a realistic workload.

Watch the segment from 17:30 to 19:19. The speaker describes four key querying patterns and their impact on GPU memory and compute. This is the foundation for creating a realistic traffic mix.

To simulate such a mix, we won't generate random data. Instead, we'll use a real-world dataset like ShareGPT, which contains a large number of diverse user conversations.

Furthermore, we need to control how requests arrive. Do they come all at once, or spaced out over time? A natural, "realistic" pattern of requests from independent users can be modeled as a Poisson process. For this, we can set a target request rate and introduce some randomness (burstiness=1.0) to mimic this process.

2. The Tool and The Target: GuideLLM and vLLM

To run our benchmark, we need a powerful tool and a worthy target for comparison.

The Target: vLLM

Throughout this course, we've often referred to vLLM. It is widely considered the state-of-the-art open-source framework for LLM inference, making it the perfect benchmark target. Its high performance is not magic; it's the result of sophisticated architectural innovations.

Accelerating LLM Inference with vLLM

This video from Databricks provides an excellent introduction to vLLM, explaining the problems it solves and its key technical innovation, PagedAttention.

Watch from 01:30 to 04:05. This will give you context on why vLLM was created and why its approach to KV cache management is so effective.

vLLM's performance comes from a combination of features like PagedAttention, FlashAttention integration, and advanced scheduling, which we will analyze later.

The Tool: GuideLLM

Instead of writing a custom benchmarking script, we'll use GuideLLM, a powerful benchmarking platform developed by the vLLM team. It is specifically designed for evaluating LLM inference systems.

vllm-project/guidellm: Evaluate and Enhance Your LLM ... - GitHub

The GuideLLM GitHub repository provides a comprehensive overview of the tool's capabilities. We'll read a few key sections to understand why it's the right choice for our task.

Please read the 'Overview', 'Why GuideLLM?', and 'Comparisons' sections on the main page. This will explain its focus on LLM-specific metrics, realistic traffic generation, and detailed reporting, making it ideal for our analysis.

3. Practical Benchmark: Step-by-Step

Now, let's put theory into practice. You will launch your engine and vLLM, then use GuideLLM to send an identical, realistic workload to both.

Step 1: Install Tools

First, ensure you have vLLM and GuideLLM installed in your Python environment.




# Install vLLM
pip install vllm




# Install GuideLLM with recommended dependencies
pip install guidellm[recommended]

Step 2: Prepare the Servers

You will run one server at a time on localhost:8000. Let's assume the small model you've been using for your custom engine is TinyLlama/TinyLlama-1.1B-Chat-v1.0.

Launch vLLM Server:
In a terminal, start the vLLM server. It will automatically download the model if it's not cached.

python -m vllm.entrypoints.openai.api_server \
  --model TinyLlama/TinyLlama-1.1B-Chat-v1.0 \
  --port 8000

Wait for the server to report that it's running.

Launch Your Custom Engine Server:
After you've benchmarked vLLM, stop it (Ctrl+C) and start your own server as you implemented in the previous lesson.

python main_engine.py # Or whatever you named your main script

Step 3: Run the Benchmark

We will use GuideLLM to send requests from the ShareGPT dataset, simulating a constant arrival rate of 2 requests per second.

Create a script or run the following command in a new terminal. You will run this command twice: once while the vLLM server is running, and once while your custom engine is running.

Benchmark Command:

guidellm benchmark run \
  --target http://localhost:8000/v1 \
  --model TinyLlama/TinyLlama-1.1B-Chat-v1.0 \
  --dataset "HuggingFaceH4/sharegpt_vicuna_unfiltered_cleaned_split" \
  --dataset-text-column "text" \
  --dataset-split "train" \
  --profile constant \
  --rate 2 \
  --max-requests 100 \
  --warmup 0.1 \
  --output-path ./results_vllm  # Change to ./results_custom_engine for the second run

Explanation of Parameters:

  • --target: The API endpoint of the server being tested.
  • --profile constant --rate 2: Sends requests at a steady rate of 2 per second (a Poisson process would use --profile poisson).
  • --max-requests 100: The test will stop after 100 successful requests.
  • --warmup 0.1: The first 10% of requests are not included in the results to allow the system to "warm up".
  • --output-path: The directory where results will be saved. Remember to change this for each run!

Execute this command for both servers and proceed once you have two result directories: results_vllm and results_custom_engine.

4. Analysis: Architectural Strengths and Weaknesses

This is the most critical part of the lesson. We will examine the results and connect them back to the architectural decisions you made when building your engine.

Step 1: Inspect the Reports

Navigate into your results directories. GuideLLM produces several files, as described in its documentation. The most immediately useful are:

  • benchmarks.html: A visual report with charts of latency and throughput. Open this in your browser.
  • benchmarks.csv: A summary table perfect for direct comparison.

Open the two CSV files and place the key results side-by-side. Your table might look something like this (numbers are hypothetical):

Metric Custom Engine vLLM
Achieved Throughput (tok/s) 450.5 780.2
Mean TTFT (ms) 310.8 185.4
Mean ITL (ms) 12.1 7.5
P99 End-to-End Latency (s) 6.2 3.8

Step 2: The "Why" - Connecting Metrics to Architecture

The data clearly shows vLLM is outperforming your custom engine. This is expected. The crucial task is to explain why.

1. Why is vLLM's throughput higher and latencies (TTFT, ITL) lower?

This is likely due to several advanced optimizations in vLLM that your engine, by design, does not have.

  • Continuous Batching vs. Chunked Prefill: Your engine implements continuous batching, which is a massive improvement over static batching. However, vLLM uses an even more advanced variant.

    Batching Strategies
    This image contrasts basic batching with continuous batching. vLLM takes this a step further.

    Watch the following clip to understand "chunked prefill," an optimization that allows vLLM to interleave prefill and decode operations, preventing decode requests from being stalled by new, long prefill requests. This directly improves throughput and reduces latency for users already in a generation loop.

Accelerating LLM Inference with vLLM

This segment of the vLLM presentation explains the concept of chunked prefill, a key optimization over naive continuous batching.

Watch the detailed explanation from 18:02 to 24:20. Pay close attention to the diagrams showing how requests can be paused in standard continuous batching, and how chunked prefill avoids this, leading to better GPU utilization and lower perceived latency.

  • Optimized Kernels (FlashAttention): vLLM integrates state-of-the-art, hardware-aware CUDA kernels like FlashAttention. As you learned in Module 5, FlashAttention minimizes slow HBM reads/writes, dramatically speeding up the attention mechanism. Your engine likely uses PyTorch's scaled dot-product attention (torch.nn.functional.scaled_dot_product_attention), which may be fast, but vLLM's deeper integration of multiple custom kernels gives it an edge in both prefill and decode.

  • Speculative Decoding: For some models and hardware, vLLM can employ speculative decoding. This advanced technique uses a small "draft" model to predict several tokens ahead, then uses the large model to verify them in a single step. This can drastically reduce the effective Inter-Token Latency (ITL). Your engine generates one token at a time, making each decode step a bottleneck.

2. What are the architectural strengths of your custom engine?

Despite the performance gap, your engine has significant architectural strengths that are highly valuable in real-world scenarios:

  • Transparency and Control: You wrote every line of the scheduler, the paged KV cache manager, and the execution loop. You understand its behavior completely. This is invaluable for debugging, profiling, and predicting performance on new workloads. With vLLM, much of the internal logic is a complex black box.
  • Extensibility: Because you own the codebase, you can easily experiment with novel ideas. Want to test a new scheduling algorithm? Or a custom KV cache eviction policy? In your engine, this is a matter of modifying a few functions. In vLLM, it would require navigating a large, complex, and highly optimized C++/CUDA/Python codebase.
  • Minimalism: Your engine is built for a specific purpose and has exactly the features you implemented. It's not bloated with support for dozens of model architectures, quantization schemes, or hardware backends you may not need. This simplicity makes it easier to maintain, deploy, and reason about.

3. How could you close the performance gap?

Based on this analysis, you can now identify a clear roadmap for improving your engine:

  • Kernel Optimization: Integrate torch.compile more deeply or even write custom Triton kernels for key operations.
  • Scheduler Enhancement: Evolve your continuous batching scheduler towards a chunked prefill model to better handle mixed workloads.
  • Advanced Features: As a major research project, you could attempt to implement a form of speculative decoding.

Conclusion

Congratulations on completing the capstone project! You have not only built a sophisticated LLM inference engine from the ground up but have now also subjected it to a rigorous, professional benchmarking process.

Key Takeaways:

  • Benchmarking is an analytical process: It's about designing realistic workloads (using real datasets and traffic patterns), measuring key metrics (TTFT, ITL, throughput), and connecting the results back to architectural design.
  • Performance comes from deep optimization: State-of-the-art frameworks like vLLM achieve high performance through a combination of advanced algorithms (PagedAttention, chunked prefill), hardware-specific kernels (FlashAttention), and complex features (speculative decoding).
  • The value of a custom engine is control: While it may not win raw performance benchmarks against a heavily funded and community-developed project, a custom engine provides unparalleled transparency, control, and extensibility, which are critical for R&D and building highly specialized AI systems.

Preview of the Next Module:
You have now mastered the art of building and analyzing the core inference engine. In the final module of the course, Production Systems: Concurrency and Stateful Services, we will shift our focus from building the engine to operating it as part of a robust, production-grade service. You will learn to handle real-world challenges like managing multi-turn conversations, implementing stateful tool-calling, and ensuring system stability with rate limiting and overload protection.

Can't find a good explanation? Sign up and we'll make it for you

Sign up