One GPU beat two.
Two GPUs ran 1.56× faster but cost 28% more per token.
Qwen2.5-72B BF16, vLLM 0.23 ROCm, 32 concurrent requests. 4 + 3 checks, Oct 1 2026.
Qwen2.5-72B on one MI300X: $1.67 per million tokens. On two: $2.14. Measured, not guessed.
Watch the 75-second filmthrottle check prints your $ per million tokens, with a 95% interval.
A flag, the model, the engine, the quantization.
Only when the change beats your measured run-to-run noise.
Two GPUs ran 1.56× faster but cost 28% more per token.
Qwen2.5-72B BF16, vLLM 0.23 ROCm, 32 concurrent requests. 4 + 3 checks, Oct 1 2026.
--max-num-seqs 8 at 32 concurrent requests.
Qwen2.5-7B, vLLM 0.23 ROCm. MORE EXPENSIVE in 4 of 4 checks.
max_num_seqs 1 → 8.
Qwen2.5-0.5B, vLLM 0.16, A100 80GB. Counterbalanced golden run.
GPU $/hr are list prices (assumed). Tokens and time are measured.Every measured result → Cost Index
throttle-pro 0.4.2 on a MacBook (GPU rate assumed at $1.50/hr), and the A100 result from the saved golden run.Measured on your stack, plus verdicts on up to 3 changes and a CI gate that fails a costlier deploy.
Book a 15-minute call →No savings guarantee. What the audit covers