LLM evaluation
A Fair LLM Inference Benchmark for vLLM and SGLang
Compare serving stacks with identical workloads, time to first token, token latency, throughput, and cache controls.
Serving benchmarks change dramatically with prompt length, output length, concurrency, and cache reuse. A fair comparison fixes the workload before naming a faster stack.
Control the setup
Use the same model weights, quantization, GPU hardware, memory budget, and output settings for each server. Record software versions and startup parameters. Build a request mix resembling production: short chat turns, long context requests, and the output lengths users actually need. Warm both systems equally, then run repeated trials.
Report a latency curve
Measure time to first token, inter-token latency, end-to-end latency, successful request rate, and output tokens per second at several offered loads. Include median and tail latency, not only peak throughput. vLLM's benchmark guide defines these measures and warns that repeated prompts can benefit from prefix-cache reuse. State whether cache hits are representative of real traffic.
Find the operating point
Plot throughput against latency and identify the highest load that still meets the product's response-time target. Note out-of-memory errors, queue growth, and output quality regressions. This guide is a method; it does not report a fresh vLLM versus SGLang winner. See our inference comparison for model and serving context.