FP8 quantization in vLLM: 47% cheaper tokens, until you read the answers
FP8 is the easiest cost cut in vLLM: one flag. We measured it on an AMD MI300X, and the biggest saving came with broken answers.
The setup
vLLM 0.23.1 on ROCm, one MI300X per server at Hot Aisle's $2.99/hr list price, 32 concurrent requests, 256 max output tokens, cold cache. Every BF16 baseline was checked four times to measure noise; every FP8 result three times.
Results
| Model | BF16 | FP8 type | FP8 $/M | Answers |
|---|---|---|---|---|
| Qwen2.5-72B | $1.67 | Pre-quantized (RedHatAI FP8-dynamic) | $1.13 (−32%) | Correct |
| Qwen2.5-7B | $0.227 | Pre-quantized | $0.186 (−18%) | Correct |
| Qwen2.5-32B | $0.76 | On the fly (--quantization fp8) | $0.45 (−41%) | Broken: "!!!!!!" |
| Qwen2.5-72B | $1.67 | On the fly | $0.87 (−47%) | Every request hit the token limit |
The trap
Asked to summarize the causes of World War I in about 80 words, the BF16 32B model wrote an 89-token answer. The on-the-fly FP8 version wrote 256 exclamation marks and stopped only because it hit the token limit. Across the whole run, every FP8 request ran to the 256-token cap: 8,192 output tokens per batch of 32 requests, against 6,883 for BF16.
That is exactly why its cost per token looked so good. A model that never stops generates tokens quickly, and cost per token only counts tokens, not whether they mean anything. The pre-quantized checkpoints produced output lengths within a few percent of BF16 (72B: 6,361 vs 6,328 tokens per batch) and answered correctly.
The 72B on-the-fly run showed the same signal, every request at the cap, though we did not read its answers before the server stopped, so we only call it unverified. These results are for this vLLM ROCm build; other versions and GPUs may behave differently, which is the point: check yours.
How to check FP8 on your own server
Measure BF16 three times, switch to FP8, measure again, and read a handful of answers from both. The cheapest check of all is output length: at temperature 0 the same prompts should get answers of about the same length.
pipx install throttle-pro
throttle check \
--url http://localhost:8000 \
--model Qwen/Qwen2.5-72B-Instruct \
--gpu-hourly-rate 2.99
The next Throttle release adds an OUTPUT CHANGED verdict: when the answers' length per request moves more than 10% or many more requests hit the token limit, it refuses to call the change CHEAPER. Replayed on the runs above, it flags both on-the-fly FP8 results and leaves the pre-quantized ones as CHEAPER.
Takeaways
- Prefer pre-quantized FP8 checkpoints over on-the-fly quantization when one exists for your model.
- A big drop in cost per token is a reason to read the answers, not to ship.
- Six spot-check prompts catch broken output; they don't measure accuracy. Run your own evals before production.
FAQ
Does FP8 quantization reduce LLM cost per token?
Usually. In our MI300X runs, pre-quantized FP8 cut cost per million output tokens by 32% on Qwen2.5-72B and 18% on Qwen2.5-7B with correct answers.
Is vLLM's on-the-fly FP8 safe to use?
Check your own build. In our vLLM 0.23.1 ROCm runs, on-the-fly --quantization fp8 produced broken output on Qwen2.5-32B while looking 41% cheaper. Pre-quantized checkpoints did not have this problem.
How do I know if quantization broke my model?
Compare a few answers at temperature 0 against the unquantized model, and watch output length per request: a jump, or every request hitting max tokens, is a red flag.
What is the difference between pre-quantized and on-the-fly FP8?
A pre-quantized checkpoint stores FP8 weights calibrated ahead of time. On-the-fly FP8 converts BF16 weights when vLLM loads the model.
Related guides
vLLM cost per token: measure it, then change one flag at a time
Measure vLLM cost per million tokens on your own GPU, then test max-num-seqs, tensor parallel size and FP8 one at a time. Real results from AMD MI300X and NVIDIA A100 runs.
CalculatorLLM cost per million tokens: the formula, a calculator, and real numbers
How to calculate LLM inference cost per million tokens from your GPU price and throughput, with a calculator, a worked example from a real AMD MI300X run, and the pitfalls that make the number wrong.