LLM evaluation

A Fair LLM Inference Benchmark for vLLM and SGLang

Compare serving stacks with identical workloads, time to first token, token latency, throughput, and cache controls.

Serving benchmarks change dramatically with prompt length, output length, concurrency, and cache reuse. A fair comparison fixes the workload before naming a faster stack.

Control the setup

Use the same model weights, quantization, GPU hardware, memory budget, and output settings for each server. Record software versions and startup parameters. Build a request mix resembling production: short chat turns, long context requests, and the output lengths users actually need. Warm both systems equally, then run repeated trials.

Report a latency curve

Measure time to first token, inter-token latency, end-to-end latency, successful request rate, and output tokens per second at several offered loads. Include median and tail latency, not only peak throughput. vLLM's benchmark guide defines these measures and warns that repeated prompts can benefit from prefix-cache reuse. State whether cache hits are representative of real traffic.

Find the operating point

Plot throughput against latency and identify the highest load that still meets the product's response-time target. Note out-of-memory errors, queue growth, and output quality regressions. This guide is a method; it does not report a fresh vLLM versus SGLang winner. See our inference comparison for model and serving context.

Sources

  1. vLLM: Benchmark CLI and latency definitions
← All articlesRelated article →