Skip to main content
grasp.study
Back to course
Log in
Get started
Someone wanted to learn this too, so Grasp built them a personal learning path.
Create your own
AI Systems Engineering for LLMs
ยท
Module 10
Evaluating Production Inference Frameworks
1
Deploy and Benchmark a 7B Model with vLLM
Deploy a 7B model using vLLM and benchmark its throughput, latency, and GPU utilization.
2
SGLang vs. vLLM: Baseline Performance Benchmark
Deploy the same model using SGLang and benchmark its baseline performance against vLLM.
3
SGLang vs. vLLM: Structured Generation Comparison
Compare the structured generation capabilities (e.g., for JSON output) of SGLang vs. vLLM.
4
Optimizing LLM Deployment with TensorRT-LLM: Benchmarking Against vLLM and SGLang
Build and deploy a model using TensorRT-LLM by compiling it into an optimized engine, then benchmark its performance against vLLM and SGLang.
5
LLM-d Architecture & Extreme Scale Use Cases
Explain the disaggregated serving architecture of systems like LLM-d and their target use cases for extreme scale.
6
VLLM, SGLang, and TensorRT-LLM: A Workload-Based Decision Matrix
Create a decision matrix for selecting between vLLM, SGLang, and TensorRT-LLM based on workload characteristics (e.g., latency vs. throughput, structured vs. open-ended generation).
7
Deploying and Load-Testing Large Language Models on Multi-GPU Systems
Deploy a 70B+ model on a multi-GPU setup using a chosen framework and load-test it with 500 concurrent users.
Previous module
Distributed Inference for Large Models
Next module
Capstone Project: Build a Custom Inference Engine