What tokens really cost.
Real serving changes on real hardware, measured in dollars per million output tokens. A change only counts as CHEAPER or MORE EXPENSIVE when it beats run-to-run noise measured from repeat checks. When the answers change shape, it's flagged OUTPUT CHANGED: a cheaper token isn't a saving if the answers broke.
| Model · engine | GPU | Change | Verdict | Details | |||
|---|---|---|---|---|---|---|---|
| Qwen2.5-72B-InstructvLLM 0.23.1 (ROCm) | 1x AMD MI300X (Hot Aisle) | BF16 → pre-quantized FP8 checkpoint (RedHatAI/Qwen2.5-72B-Instruct-FP8-dynamic) | $1.67$1.67–$1.68 | $1.13$1.13–$1.13 | -32.5% | CHEAPER3 of 3 | |
FP8 server ran on the VM's second MI300X, the BF16 baseline on the first (same GPU model; Throttle flagged the different endpoint).
Every check of this change
## Throttle check: Qwen/Qwen2.5-72B-Instruct | | | |---|---| | Throttle version | 0.5.1 | | Engine | vllm-rocm (from --config) | | GPU | MI300X | | Model | Qwen/Qwen2.5-72B-Instruct | | GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) | | Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) | | Prompt cache | cold (unique prompt per request) | | Label | fp8-prequant | | Check | 20261001T063943Z-bcd9daf2 (2026-10-01) | **Result:** $1.13 per million output tokens (95% CI $1.13 to $1.13, 5 blocks) [MEASURED] **Compared with** check 20261001T062110Z-87304e73: - before: $1.67/M output tokens (95% CI $1.67 to $1.68) - after: $1.13/M output tokens (95% CI $1.13 to $1.13) - change: -32.4% - noise floor: calibrated: bound 0.6% (t x sqrt(2) x 0.2% run-to-run SD, 4 df) - **Verdict: CHEAPER**, the -32.4% change is larger than the run-to-run noise bound (0.6% = t x sqrt(2) x 0.2% run-to-run SD, 4 df) and the 95% confidence intervals do not overlap **What changed in config:** - endpoint: changed (URLs are not shared) - `config.quantization`: none -> fp8-redhat-dynamic - `label`: baseline -> fp8-prequant - `server.[host removed]`: Qwen/Qwen2.5-72B-Instruct -> RedHatAI/Qwen2.5-72B-Instruct-FP8-dynamic - `vllm.cache_config.num_gpu_blocks`: 7776 -> 21116 _Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._ | |||||||
| Qwen2.5-72B-InstructvLLM 0.23.1 (ROCm) | 1x AMD MI300X (Hot Aisle) | BF16 → vLLM on-the-fly FP8 (--quantization fp8) |
$1.67$1.67–$1.68 | $0.872$0.868–$0.876 | -47.9 to -47.0% | OUTPUT CHANGED3 of 3 | |
Not a saving. Looked 47% cheaper per token; excluded as a saving because the answers changed shape.
Every check of this change
## Throttle check: Qwen/Qwen2.5-72B-Instruct | | | |---|---| | Throttle version | 0.5.1 | | Engine | vllm-rocm (from --config) | | GPU | MI300X | | Model | Qwen/Qwen2.5-72B-Instruct | | GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) | | Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) | | Prompt cache | cold (unique prompt per request) | | Label | fp8 | | Check | 20261001T062654Z-8d20c47c (2026-10-01) | **Result:** $0.8721 per million output tokens (95% CI $0.8682 to $0.8759, 5 blocks) [MEASURED] **Compared with** check 20261001T062110Z-87304e73: - before: $1.67/M output tokens (95% CI $1.67 to $1.68) - after: $0.8721/M output tokens (95% CI $0.8682 to $0.8759) - change: -47.9% - noise floor: calibrated: bound 2.5% (t x sqrt(2) x 0.6% run-to-run SD, 4 df) - **Verdict: OUTPUT CHANGED**, the server's answers changed: output length per request moved +29.5% (197.7 -> 256.0 tokens; more than 10%). At temperature 0 the same prompts should get the same answers, so this $/M difference may come from broken or padded output, not a faster server. Read a few answers from both configs before trusting it **What changed in config:** - `config.quantization`: none -> fp8 - `label`: baseline -> fp8 - `vllm.cache_config.num_gpu_blocks`: 7776 -> 21166 _Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._ | |||||||
| Qwen2.5-32B-InstructvLLM 0.23.1 (ROCm) | 1x AMD MI300X (Hot Aisle) | BF16 → vLLM on-the-fly FP8 (--quantization fp8) |
$0.761$0.751–$0.771 | $0.450$0.447–$0.454 | -41.1 to -40.8% | OUTPUT CHANGED3 of 3 | |
Not a saving. Looked 41% cheaper per token; the answers were broken.
Every check of this change
## Throttle check: Qwen/Qwen2.5-32B-Instruct | | | |---|---| | Throttle version | 0.5.1 | | Engine | vllm-rocm (from --config) | | GPU | MI300X | | Model | Qwen/Qwen2.5-32B-Instruct | | GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) | | Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) | | Prompt cache | cold (unique prompt per request) | | Label | fp8 | | Check | 20261001T062205Z-cfa31769 (2026-10-01) | **Result:** $0.4503 per million output tokens (95% CI $0.4466 to $0.4541, 5 blocks) [MEASURED] **Compared with** check 20261001T061821Z-35ff548e: - before: $0.7608/M output tokens (95% CI $0.7508 to $0.7709) - after: $0.4503/M output tokens (95% CI $0.4466 to $0.4541) - change: -40.8% - noise floor: calibrated: bound 4.9% (t x sqrt(2) x 1.2% run-to-run SD, 4 df) - **Verdict: OUTPUT CHANGED**, the server's answers changed: output length per request moved +17.7% (217.5 -> 256.0 tokens; more than 10%). At temperature 0 the same prompts should get the same answers, so this $/M difference may come from broken or padded output, not a faster server. Read a few answers from both configs before trusting it **What changed in config:** - `config.quantization`: none -> fp8 - `label`: baseline -> fp8 - `vllm.cache_config.num_gpu_blocks`: 28821 -> 36228 _Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._ | |||||||
| Qwen2.5-7B-InstructvLLM 0.23.1 (ROCm) | 1x AMD MI300X (Hot Aisle) | BF16 → pre-quantized FP8 checkpoint (RedHatAI/Qwen2.5-7B-Instruct-FP8-dynamic) | $0.227$0.222–$0.233 | $0.186$0.182–$0.190 | -18.8 to -18.3% | CHEAPER3 of 3 | |
Every check of this change
## Throttle check: Qwen/Qwen2.5-7B-Instruct | | | |---|---| | Throttle version | 0.5.1 | | Engine | vllm-rocm (from --config) | | GPU | MI300X | | Model | Qwen/Qwen2.5-7B-Instruct | | GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) | | Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) | | Prompt cache | cold (unique prompt per request) | | Label | fp8-prequant | | Check | 20261001T065652Z-411315d1 (2026-10-01) | **Result:** $0.1856 per million output tokens (95% CI $0.1817 to $0.1895, 5 blocks) [MEASURED] **Compared with** check 20261001T054905Z-0f15c02b: - before: $0.2272/M output tokens (95% CI $0.2217 to $0.2327) - after: $0.1856/M output tokens (95% CI $0.1817 to $0.1895) - change: -18.3% - noise floor: calibrated: bound 2.3% (t x sqrt(2) x 0.6% run-to-run SD, 4 df) - **Verdict: CHEAPER**, the -18.3% change is larger than the run-to-run noise bound (2.3% = t x sqrt(2) x 0.6% run-to-run SD, 4 df) and the 95% confidence intervals do not overlap **What changed in config:** - `config.max_num_seqs`: default -> (not set) - `config.quantization`: (not set) -> fp8-redhat-dynamic - `label`: baseline -> fp8-prequant - `server.[host removed]`: Qwen/Qwen2.5-7B-Instruct -> RedHatAI/Qwen2.5-7B-Instruct-FP8-dynamic - `vllm.cache_config.num_gpu_blocks`: 186542 -> 193570 _Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._ | |||||||
| Qwen2.5-7B-InstructvLLM 0.23.1 (ROCm) | 1x AMD MI300X (Hot Aisle) | BF16 → vLLM on-the-fly FP8 (--quantization fp8) |
$0.227$0.222–$0.233 | $0.186$0.169–$0.203 | -18.1 to -15.5% | CHEAPER2 of 2 | |
Every check of this change
## Throttle check: Qwen/Qwen2.5-7B-Instruct | | | |---|---| | Throttle version | 0.5.1 | | Engine | vllm-rocm (from --config) | | GPU | MI300X | | Model | Qwen/Qwen2.5-7B-Instruct | | GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) | | Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) | | Prompt cache | cold (unique prompt per request) | | Label | fp8 | | Check | 20261001T055349Z-ec722b14 (2026-10-01) | **Result:** $0.1860 per million output tokens (95% CI $0.1691 to $0.2030, 5 blocks) [MEASURED] **Compared with** check 20261001T054905Z-0f15c02b: - before: $0.2272/M output tokens (95% CI $0.2217 to $0.2327) - after: $0.1860/M output tokens (95% CI $0.1691 to $0.2030) - change: -18.1% - noise floor: calibrated: bound 2.8% (t x sqrt(2) x 0.6% run-to-run SD, 3 df) - **Verdict: CHEAPER**, the -18.1% change is larger than the run-to-run noise bound (2.8% = t x sqrt(2) x 0.6% run-to-run SD, 3 df) and the 95% confidence intervals do not overlap **What changed in config:** - `config.quantization`: (not set) -> fp8 - `label`: baseline -> fp8 - `vllm.cache_config.num_gpu_blocks`: 186542 -> 193542 _Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._ | |||||||
| Qwen2.5-7B-InstructvLLM 0.23.1 (ROCm) | 1x AMD MI300X (Hot Aisle) | vLLM defaults → --max-num-seqs 8 (at 32 concurrent requests) |
$0.227$0.222–$0.233 | $0.700$0.684–$0.716 | +208.0 to +213.4% | MORE EXPENSIVE4 of 4 | |
A deliberately bad setting for this load: shows what a harmless-looking flag costs.
Every check of this change
## Throttle check: Qwen/Qwen2.5-7B-Instruct | | | |---|---| | Throttle version | 0.5.1 | | Engine | vllm-rocm (from --config) | | GPU | MI300X | | Model | Qwen/Qwen2.5-7B-Instruct | | GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) | | Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) | | Prompt cache | cold (unique prompt per request) | | Label | seqs8 | | Check | 20261001T055720Z-7973bb09 (2026-10-01) | **Result:** $0.6998 per million output tokens (95% CI $0.6839 to $0.7158, 5 blocks) [MEASURED] **Compared with** check 20261001T054905Z-0f15c02b: - before: $0.2272/M output tokens (95% CI $0.2217 to $0.2327) - after: $0.6998/M output tokens (95% CI $0.6839 to $0.7158) - change: +208.0% - noise floor: calibrated: bound 2.8% (t x sqrt(2) x 0.8% run-to-run SD, 5 df) - **Verdict: MORE EXPENSIVE**, the +208.0% change is larger than the run-to-run noise bound (2.8% = t x sqrt(2) x 0.8% run-to-run SD, 5 df) and the 95% confidence intervals do not overlap **What changed in config:** - `config.max_num_seqs`: default -> 8 - `label`: baseline -> seqs8 - `vllm.cache_config.num_gpu_blocks`: 186542 -> 187972 _Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._ | |||||||
| Qwen2.5-72B-InstructvLLM 0.23.1 (ROCm) | 1x vs 2x AMD MI300X (Hot Aisle) | 1 GPU (TP1, $2.99/hr) → 2 GPUs (--tensor-parallel-size 2, $5.98/hr), same BF16 weights |
$1.66$1.66–$1.67 | $2.13$2.13–$2.14 | +28.1 to +28.8% | MORE EXPENSIVE3 of 3 | |
Throttle 0.5.1 prints CHEAPER here because its verdict rescales to the baseline's GPU rate; this row judges real dollars with the same rule (2 GPUs: 1.56x the speed, 2x the price).
Every check of this change
## Throttle check: Qwen/Qwen2.5-72B-Instruct | | | |---|---| | Throttle version | 0.5.1 | | Engine | vllm (from --config) | | GPU | MI300X | | Model | Qwen/Qwen2.5-72B-Instruct | | GPU hourly rate | $5.98/hr (ASSUMED, user-supplied) | | Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) | | Prompt cache | cold (unique prompt per request) | | Label | tp2 | | Check | 20261001T192700Z-d565005e (2026-10-01) | **Result:** $2.13 per million output tokens (95% CI $2.13 to $2.14, 5 blocks) [MEASURED] **Compared with** check 20261001T192145Z-38a09569: - before: $1.66/M output tokens (95% CI $1.66 to $1.67) - after: $2.13/M output tokens (95% CI $2.13 to $2.14) - change: +28.1% at the GPU rates as typed - measured change at the baseline's GPU rate: +28.1% - noise floor: calibrated: bound 0.8% (t x sqrt(2) x 0.2% run-to-run SD, 4 df) - **Verdict: MORE EXPENSIVE**, in real dollars (GPU count and price changed) the +28.1% change is larger than the run-to-run noise bound (0.8%) and the 95% confidence intervals do not overlap **What changed in config:** - `config.gpus`: 1 -> 2 - `config.tensor_parallel`: 1 -> 2 - `gpu_hourly_rate_usd (ASSUMED)`: 2.9900 -> 5.9800 - `label`: tp1 -> tp2 - `vllm.cache_config.num_gpu_blocks`: 7776 -> 43167 _Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._ | |||||||
| llama3.2:1bOllama 0.32.5 | Apple M3 Pro (MacBook) | OLLAMA_NUM_PARALLEL 1 → 4 | $5.93$5.80–$6.06 | $2.01$1.96–$2.06 | -66.1 to -61.7% | CHEAPER2 of 2 | |
Laptop; the $/hr is an assumed rate for illustration.
Every check of this change
## Throttle check: llama3.2:1b | | | |---|---| | Throttle version | 0.5.1 | | Engine | Ollama (likely; /v1/models owned_by=library) | | GPU | not given (add --config gpu=NAME) | | Model | llama3.2:1b | | GPU hourly rate | $1.50/hr (ASSUMED, user-supplied) | | Workload | 5 blocks x 8 requests, concurrency 4, max 128 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) | | Prompt cache | cold (unique prompt per request) | | Label | parallel-4 | | Check | 20260930T064345Z-bef600a8 (2026-09-30) | **Result:** $2.01 per million output tokens (95% CI $1.96 to $2.06, 5 blocks) [MEASURED] **Compared with** check 20260930T064223Z-93b28975: - before: $5.93/M output tokens (95% CI $5.80 to $6.06) - after: $2.01/M output tokens (95% CI $1.96 to $2.06) - change: -66.1% - noise floor: calibrated: bound 15.0% (t x sqrt(2) x 3.3% run-to-run SD, 3 df) - **Verdict: CHEAPER**, the -66.1% change is larger than the run-to-run noise bound (15.0% = t x sqrt(2) x 3.3% run-to-run SD, 3 df) and the 95% confidence intervals do not overlap **What changed in config:** - `config.ollama_num_parallel`: 1 -> 4 - `label`: parallel-1 -> parallel-4 _Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._ | |||||||
| llama3.2:3bOllama 0.32.5 | Apple M3 Pro (MacBook) | Q4_K_M → Q8_0 weights (OLLAMA_NUM_PARALLEL=4) | $5.18$5.08–$5.27 | $4.70$4.53–$4.87 | -10.3 to -9.2% | CHEAPER3 of 3 | |
Laptop; the $/hr is an assumed rate. Q4/Q8 order was interleaved to rule out drift.
Every check of this change
## Throttle check: llama3.2:3b-instruct-q8_0 | | | |---|---| | Throttle version | 0.5.1 | | Engine | Ollama (likely; /v1/models owned_by=library) | | GPU | not given (add --config gpu=NAME) | | Model | llama3.2:3b-instruct-q8_0 | | GPU hourly rate | $1.50/hr (ASSUMED, user-supplied) | | Workload | 5 blocks x 8 requests, concurrency 4, max 128 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) | | Prompt cache | cold (unique prompt per request) | | Label | q8 | | Check | 20261001T035022Z-29452cbf (2026-10-01) | **Result:** $4.70 per million output tokens (95% CI $4.53 to $4.87, 5 blocks) [MEASURED] **Compared with** check 20261001T034456Z-06d7816a: - before: $5.18/M output tokens (95% CI $5.08 to $5.27) - after: $4.70/M output tokens (95% CI $4.53 to $4.87) - change: -9.2% - noise floor: calibrated: bound 3.9% (t x sqrt(2) x 1.1% run-to-run SD, 6 df) - **Verdict: CHEAPER**, the -9.2% change is larger than the run-to-run noise bound (3.9% = t x sqrt(2) x 1.1% run-to-run SD, 6 df) and the 95% confidence intervals do not overlap **What changed in config:** - `config.quant`: Q4_K_M -> Q8_0 - `label`: q4 -> q8 - `request.model`: llama3.2:3b -> llama3.2:3b-instruct-q8_0 _Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._ | |||||||
| llama3.2:3b → llama3.2:1bOllama | Apple M3 Pro (MacBook) | Swap llama3.2:3b for llama3.2:1b | $8.63$8.27–$9.00 | $5.91$4.85–$6.97 | -35.9 to -31.6% | CHEAPER2 of 2 | |
Laptop; the $/hr is an assumed rate.
Every check of this change
## Throttle check: llama3.2:1b | | | |---|---| | Throttle version | 0.4.2 | | Engine | Ollama (likely; /v1/models owned_by=library) | | GPU | apple-m3-pro | | Model | llama3.2:1b | | GPU hourly rate | $1.50/hr (ASSUMED, user-supplied) | | Workload | 3 blocks x 4 requests, concurrency 2, max 64 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) | | Prompt cache | cold (unique prompt per request) | | Label | smaller-1b | | Check | 20260927T194337Z-41f5e1f0 (2026-09-27) | **Result:** $5.91 per million output tokens (95% CI $4.85 to $6.97, 3 blocks) [MEASURED] **Compared with** check 20260927T194311Z-48d7dd1d: - before: $8.63/M output tokens (95% CI $8.27 to $9.00) - after: $5.91/M output tokens (95% CI $4.85 to $6.97) - change: -31.6% - noise floor: calibrated: bound 5.0% (t x sqrt(2) x 1.1% run-to-run SD, 3 df) - **Verdict: CHEAPER**, the -31.6% change is larger than the run-to-run noise bound (5.0% = t x sqrt(2) x 1.1% run-to-run SD, 3 df) and the 95% confidence intervals do not overlap **What changed in config:** - `label`: baseline-3b -> smaller-1b - `request.model`: llama3.2:3b -> llama3.2:1b _Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._ | |||||||
| Qwen2.5-0.5B-InstructvLLM 0.16.0 | 1x NVIDIA A100 80GB PCIe (RunPod) | max_num_seqs 1 → 8 | $0.746 | $0.234 | -68.6% | CHEAPERdecision-eligible (golden protocol) | |
A deliberately bad baseline (one sequence at a time) on a small model; it shows the measurement, not a saving to expect.
| |||||||
$ per million output tokens = measured wall-clock time × the GPU's hourly price ÷ measured output tokens. Ranges show every check of a change; CIs are 95%. One workload per row: your prompts, concurrency and output lengths will give different numbers. Download the data (JSON).
Add your result
The Index grows from real checks. If you serve a model on your own or rented GPUs, measure one change and send it in:
- Install Throttle:
pipx install throttle-pro - Run
throttle check3 times on your current config to measure your noise, then once after your change. - Read a few answers from both configs (the FP8 rows above are why).
- Run
throttle check --share-id <id>and paste the summary into a new Index submission. The summary leaves out URLs, hostnames and anything that looks like a secret.
Every accepted row names its GPU, engine, workload and the $/hr it assumed. Nothing is added without the check output behind it.
You pick, we measure
Every week the community votes on one serving question, and we rent the GPU, run it with Throttle and add the result here, including when the answer is "no winner." Want a question measured? Open a request on GitHub or email kushthrottle@gmail.com.