All guides Guide

vLLM cost per token: measure it, then change one flag at a time

vLLM gives you dozens of flags. Each one moves your cost per token, sometimes by 3×. Measure first, change one thing, measure again.

Step 1: a baseline with vLLM defaults

Start vLLM the way you run it in production and leave everything else at its defaults:

vllm serve Qwen/Qwen2.5-7B-Instruct --port 8000

Then measure it three times without changing anything. The repeats are not wasted: they measure your run-to-run noise, which every later comparison is judged against.

pipx install throttle-pro

throttle check \
  --url http://localhost:8000 \
  --model Qwen/Qwen2.5-7B-Instruct \
  --gpu-hourly-rate 2.99

On one AMD MI300X with vLLM 0.23.1 (ROCm), 32 concurrent requests and 256 max output tokens, four baseline checks came back at $0.227, $0.230, $0.228 and $0.227 per million output tokens: under 1% spread, which gave a noise bound of 2.8%.

Step 2: change one thing

Restart vLLM with exactly one flag changed, label the check, and compare it with the baseline:

vllm serve Qwen/Qwen2.5-7B-Instruct --port 8000 --max-num-seqs 8

throttle check --url http://localhost:8000 \
  --model Qwen/Qwen2.5-7B-Instruct --gpu-hourly-rate 2.99 \
  --label seqs8 --config max_num_seqs=8

Throttle prints the before and after, both confidence intervals, what changed in the server config, and one verdict: CHEAPER, MORE EXPENSIVE, NO WINNER or NOT CALIBRATED.

What real flag changes did

ChangeHardwareBeforeAfterVerdict
--max-num-seqs 8 at 32 concurrent requests1× MI300X, 7B$0.227/M$0.700/MMORE EXPENSIVE, +208%
max_num_seqs 1 → 8 (deliberately bad baseline)1× A100, 0.5B$0.746/M$0.234/MCHEAPER, −68.6%
Pre-quantized FP8 checkpoint1× MI300X, 72B$1.67/M$1.13/MCHEAPER, −32%
Tensor parallel 1 → 2MI300X, 72B$1.67/M$2.14/M+28% per token

The same flag can help or hurt. max_num_seqs caps how many sequences vLLM batches: raising it from 1 cut cost by two thirds on the A100, while lowering the default to 8 under 32 concurrent requests tripled it on the MI300X. That is why the number has to be measured on your workload, not copied from a blog. Every row has raw data on the Cost Index.

Flags worth testing

  • --max-num-seqs and --max-num-batched-tokens: batching limits. Too low leaves the GPU idle under load.
  • --tensor-parallel-size: more GPUs buy latency, not cheaper tokens. See tensor parallel cost.
  • Quantization: often the biggest win, and the biggest risk to answer quality. See FP8 cost vs quality.
  • --enable-prefix-caching: only helps if your real prompts share prefixes. Measure it with --warm-cache if they do, and cold if they don't.

Step 3: keep checking

vLLM upgrades, model swaps and config drift all move cost silently. throttle check --fail-if-costlier 5 exits non-zero when a change is calibrated MORE EXPENSIVE by 5% or more, so it can gate a deploy in CI.

FAQ

How do I measure vLLM cost per million tokens?

Run vLLM, then point throttle check at its OpenAI-compatible endpoint with your GPU's hourly price. It sends a fixed workload, measures aggregate output tokens per second, and reports $ per million tokens with a 95% confidence interval.

Which vLLM flags affect cost per token the most?

Batching limits (max-num-seqs, max-num-batched-tokens), tensor parallel size and quantization had the largest effects in our runs, from −68.6% to +208%.

Does a higher max-num-seqs always reduce cost?

Not always. It lets vLLM batch more requests, which helps under load, but the effect depends on your concurrency and memory. Measure it.

Is the vLLM default config the cheapest?

Often close, not guaranteed. In our MI300X run, lowering max-num-seqs from the default to 8 tripled cost per token at 32 concurrent requests.

Related guides