The Throttle Index · 11 changes · updated

What tokens really cost.

Real serving changes on real hardware, measured in dollars per million output tokens. A change only counts as CHEAPER or MORE EXPENSIVE when it beats run-to-run noise measured from repeat checks. When the answers change shape, it's flagged OUTPUT CHANGED: a cheaper token isn't a saving if the answers broke.

MEASURED tokens and wall-clock time ASSUMED the GPU's $/hr (a list price, stated on every row)
Model · engine GPU Change Verdict Details
Qwen2.5-72B-InstructvLLM 0.23.1 (ROCm) 1x AMD MI300X (Hot Aisle) BF16 → pre-quantized FP8 checkpoint (RedHatAI/Qwen2.5-72B-Instruct-FP8-dynamic) $1.67$1.67–$1.68 $1.13$1.13–$1.13 -32.5% CHEAPER3 of 3

FP8 server ran on the VM's second MI300X, the BF16 baseline on the first (same GPU model; Throttle flagged the different endpoint).

Answers
6 of 6 sample answers correct and coherent (temperature 0); output length per request within 1% of BF16.
GPU price
$2.99/hr ASSUMED Hot Aisle on-demand list price per MI300X GPU. Tokens and time are MEASURED.
Workload
5 blocks × 32 requests, concurrency 32, max 256 output tokens, temperature 0, cache: cold
Baseline repeats
$1.67, $1.67, $1.67, $1.67 (noise bound 0.6%)
Method
throttle check (throttle-pro 0.5.1), re-judged with the OUTPUT CHANGED rule. Date: 2026-10-01.
Source
the saved check history (summary below is Throttle's own sanitized --share-id output)

Every check of this change

  • 20261001T063815Z-8ece234d $1.13 ($1.13–$1.13) vs $1.67: CHEAPER -32.5%, noise bound 0.8%, output length +0.7%
  • 20261001T063859Z-c8674c58 $1.13 ($1.13–$1.13) vs $1.67: CHEAPER -32.4%, noise bound 0.8%, output length +0.8%
  • 20261001T063943Z-bcd9daf2 $1.13 ($1.13–$1.13) vs $1.67: CHEAPER -32.4%, noise bound 0.6%, output length +0.4%
## Throttle check: Qwen/Qwen2.5-72B-Instruct

| | |
|---|---|
| Throttle version | 0.5.1 |
| Engine | vllm-rocm (from --config) |
| GPU | MI300X |
| Model | Qwen/Qwen2.5-72B-Instruct |
| GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) |
| Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) |
| Prompt cache | cold (unique prompt per request) |
| Label | fp8-prequant |
| Check | 20261001T063943Z-bcd9daf2 (2026-10-01) |

**Result:** $1.13 per million output tokens (95% CI $1.13 to $1.13, 5 blocks) [MEASURED]

**Compared with** check 20261001T062110Z-87304e73:

- before: $1.67/M output tokens (95% CI $1.67 to $1.68)
- after: $1.13/M output tokens (95% CI $1.13 to $1.13)
- change: -32.4%
- noise floor: calibrated: bound 0.6% (t x sqrt(2) x 0.2% run-to-run SD, 4 df)
- **Verdict: CHEAPER**, the -32.4% change is larger than the run-to-run noise bound (0.6% = t x sqrt(2) x 0.2% run-to-run SD, 4 df) and the 95% confidence intervals do not overlap

**What changed in config:**

- endpoint: changed (URLs are not shared)
- `config.quantization`: none -> fp8-redhat-dynamic
- `label`: baseline -> fp8-prequant
- `server.[host removed]`: Qwen/Qwen2.5-72B-Instruct -> RedHatAI/Qwen2.5-72B-Instruct-FP8-dynamic
- `vllm.cache_config.num_gpu_blocks`: 7776 -> 21116

_Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._
Qwen2.5-72B-InstructvLLM 0.23.1 (ROCm) 1x AMD MI300X (Hot Aisle) BF16 → vLLM on-the-fly FP8 (--quantization fp8) $1.67$1.67–$1.68 $0.872$0.868–$0.876 -47.9 to -47.0% OUTPUT CHANGED3 of 3

Not a saving. Looked 47% cheaper per token; excluded as a saving because the answers changed shape.

Answers
Not sampled. Every request ran to the 256-token cap (8,016–8,192 output tokens per 32-request block vs 6,328 for BF16).
GPU price
$2.99/hr ASSUMED Hot Aisle on-demand list price per MI300X GPU. Tokens and time are MEASURED.
Workload
5 blocks × 32 requests, concurrency 32, max 256 output tokens, temperature 0, cache: cold
Baseline repeats
$1.67, $1.67, $1.67, $1.67 (noise bound 2.5%)
Method
throttle check (throttle-pro 0.5.1), re-judged with the OUTPUT CHANGED rule. Date: 2026-10-01.
Source
the saved check history (summary below is Throttle's own sanitized --share-id output)

Every check of this change

  • 20261001T062514Z-0df61c4e $0.887 ($0.859–$0.915) vs $1.67: OUTPUT CHANGED -47.0%, noise bound 0.8%, output length +28.3%
  • 20261001T062604Z-47da7f94 $0.872 ($0.869–$0.875) vs $1.67: OUTPUT CHANGED -47.9%, noise bound 0.8%, output length +29.5%
  • 20261001T062654Z-8d20c47c $0.872 ($0.868–$0.876) vs $1.67: OUTPUT CHANGED -47.9%, noise bound 2.5%, output length +29.5%
## Throttle check: Qwen/Qwen2.5-72B-Instruct

| | |
|---|---|
| Throttle version | 0.5.1 |
| Engine | vllm-rocm (from --config) |
| GPU | MI300X |
| Model | Qwen/Qwen2.5-72B-Instruct |
| GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) |
| Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) |
| Prompt cache | cold (unique prompt per request) |
| Label | fp8 |
| Check | 20261001T062654Z-8d20c47c (2026-10-01) |

**Result:** $0.8721 per million output tokens (95% CI $0.8682 to $0.8759, 5 blocks) [MEASURED]

**Compared with** check 20261001T062110Z-87304e73:

- before: $1.67/M output tokens (95% CI $1.67 to $1.68)
- after: $0.8721/M output tokens (95% CI $0.8682 to $0.8759)
- change: -47.9%
- noise floor: calibrated: bound 2.5% (t x sqrt(2) x 0.6% run-to-run SD, 4 df)
- **Verdict: OUTPUT CHANGED**, the server's answers changed: output length per request moved +29.5% (197.7 -> 256.0 tokens; more than 10%). At temperature 0 the same prompts should get the same answers, so this $/M difference may come from broken or padded output, not a faster server. Read a few answers from both configs before trusting it

**What changed in config:**

- `config.quantization`: none -> fp8
- `label`: baseline -> fp8
- `vllm.cache_config.num_gpu_blocks`: 7776 -> 21166

_Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._
Qwen2.5-32B-InstructvLLM 0.23.1 (ROCm) 1x AMD MI300X (Hot Aisle) BF16 → vLLM on-the-fly FP8 (--quantization fp8) $0.761$0.751–$0.771 $0.450$0.447–$0.454 -41.1 to -40.8% OUTPUT CHANGED3 of 3

Not a saving. Looked 41% cheaper per token; the answers were broken.

Answers
Broken: asked to summarize WWI, it answered with 256 tokens of "!!!!". Every request hit the token cap (8,192 vs 6,883 output tokens per block).
GPU price
$2.99/hr ASSUMED Hot Aisle on-demand list price per MI300X GPU. Tokens and time are MEASURED.
Workload
5 blocks × 32 requests, concurrency 32, max 256 output tokens, temperature 0, cache: cold
Baseline repeats
$0.760, $0.776, $0.782, $0.761 (noise bound 4.9%)
Method
throttle check (throttle-pro 0.5.1), re-judged with the OUTPUT CHANGED rule. Date: 2026-10-01.
Source
the saved check history (summary below is Throttle's own sanitized --share-id output)

Every check of this change

  • 20261001T062119Z-0df04e1d $0.448 ($0.445–$0.451) vs $0.761: OUTPUT CHANGED -41.1%, noise bound 6.4%, output length +17.7%
  • 20261001T062142Z-444dc655 $0.449 ($0.446–$0.452) vs $0.761: OUTPUT CHANGED -41.0%, noise bound 6.4%, output length +17.7%
  • 20261001T062205Z-cfa31769 $0.450 ($0.447–$0.454) vs $0.761: OUTPUT CHANGED -40.8%, noise bound 4.9%, output length +17.7%
## Throttle check: Qwen/Qwen2.5-32B-Instruct

| | |
|---|---|
| Throttle version | 0.5.1 |
| Engine | vllm-rocm (from --config) |
| GPU | MI300X |
| Model | Qwen/Qwen2.5-32B-Instruct |
| GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) |
| Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) |
| Prompt cache | cold (unique prompt per request) |
| Label | fp8 |
| Check | 20261001T062205Z-cfa31769 (2026-10-01) |

**Result:** $0.4503 per million output tokens (95% CI $0.4466 to $0.4541, 5 blocks) [MEASURED]

**Compared with** check 20261001T061821Z-35ff548e:

- before: $0.7608/M output tokens (95% CI $0.7508 to $0.7709)
- after: $0.4503/M output tokens (95% CI $0.4466 to $0.4541)
- change: -40.8%
- noise floor: calibrated: bound 4.9% (t x sqrt(2) x 1.2% run-to-run SD, 4 df)
- **Verdict: OUTPUT CHANGED**, the server's answers changed: output length per request moved +17.7% (217.5 -> 256.0 tokens; more than 10%). At temperature 0 the same prompts should get the same answers, so this $/M difference may come from broken or padded output, not a faster server. Read a few answers from both configs before trusting it

**What changed in config:**

- `config.quantization`: none -> fp8
- `label`: baseline -> fp8
- `vllm.cache_config.num_gpu_blocks`: 28821 -> 36228

_Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._
Qwen2.5-7B-InstructvLLM 0.23.1 (ROCm) 1x AMD MI300X (Hot Aisle) BF16 → pre-quantized FP8 checkpoint (RedHatAI/Qwen2.5-7B-Instruct-FP8-dynamic) $0.227$0.222–$0.233 $0.186$0.182–$0.190 -18.8 to -18.3% CHEAPER3 of 3
Answers
6 of 6 sample answers correct and coherent (temperature 0).
GPU price
$2.99/hr ASSUMED Hot Aisle on-demand list price per MI300X GPU. Tokens and time are MEASURED.
Workload
5 blocks × 32 requests, concurrency 32, max 256 output tokens, temperature 0, cache: cold
Baseline repeats
$0.227, $0.230, $0.228, $0.227 (noise bound 2.3%)
Method
throttle check (throttle-pro 0.5.1), re-judged with the OUTPUT CHANGED rule. Date: 2026-10-01.
Source
the saved check history (summary below is Throttle's own sanitized --share-id output)

Every check of this change

  • 20261001T065636Z-39619c63 $0.184 ($0.178–$0.190) vs $0.227: CHEAPER -18.8%, noise bound 2.8%, output length +2.2%
  • 20261001T065644Z-79de36b0 $0.186 ($0.180–$0.192) vs $0.227: CHEAPER -18.3%, noise bound 2.8%, output length +2.0%
  • 20261001T065652Z-411315d1 $0.186 ($0.182–$0.190) vs $0.227: CHEAPER -18.3%, noise bound 2.3%, output length +2.0%
## Throttle check: Qwen/Qwen2.5-7B-Instruct

| | |
|---|---|
| Throttle version | 0.5.1 |
| Engine | vllm-rocm (from --config) |
| GPU | MI300X |
| Model | Qwen/Qwen2.5-7B-Instruct |
| GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) |
| Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) |
| Prompt cache | cold (unique prompt per request) |
| Label | fp8-prequant |
| Check | 20261001T065652Z-411315d1 (2026-10-01) |

**Result:** $0.1856 per million output tokens (95% CI $0.1817 to $0.1895, 5 blocks) [MEASURED]

**Compared with** check 20261001T054905Z-0f15c02b:

- before: $0.2272/M output tokens (95% CI $0.2217 to $0.2327)
- after: $0.1856/M output tokens (95% CI $0.1817 to $0.1895)
- change: -18.3%
- noise floor: calibrated: bound 2.3% (t x sqrt(2) x 0.6% run-to-run SD, 4 df)
- **Verdict: CHEAPER**, the -18.3% change is larger than the run-to-run noise bound (2.3% = t x sqrt(2) x 0.6% run-to-run SD, 4 df) and the 95% confidence intervals do not overlap

**What changed in config:**

- `config.max_num_seqs`: default -> (not set)
- `config.quantization`: (not set) -> fp8-redhat-dynamic
- `label`: baseline -> fp8-prequant
- `server.[host removed]`: Qwen/Qwen2.5-7B-Instruct -> RedHatAI/Qwen2.5-7B-Instruct-FP8-dynamic
- `vllm.cache_config.num_gpu_blocks`: 186542 -> 193570

_Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._
Qwen2.5-7B-InstructvLLM 0.23.1 (ROCm) 1x AMD MI300X (Hot Aisle) BF16 → vLLM on-the-fly FP8 (--quantization fp8) $0.227$0.222–$0.233 $0.186$0.169–$0.203 -18.1 to -15.5% CHEAPER2 of 2
Answers
6 of 6 sample answers correct and coherent (temperature 0); output length +6–9% (under the 10% OUTPUT CHANGED line).
GPU price
$2.99/hr ASSUMED Hot Aisle on-demand list price per MI300X GPU. Tokens and time are MEASURED.
Workload
5 blocks × 32 requests, concurrency 32, max 256 output tokens, temperature 0, cache: cold
Baseline repeats
$0.227, $0.230, $0.228, $0.227 (noise bound 2.8%)
Method
throttle check (throttle-pro 0.5.1), re-judged with the OUTPUT CHANGED rule. Date: 2026-10-01.
Source
the saved check history (summary below is Throttle's own sanitized --share-id output)

Every check of this change

  • 20261001T055339Z-d24e1ac8 $0.192 ($0.180–$0.204) vs $0.227: CHEAPER -15.5%, noise bound 2.8%, output length +6.3%
  • 20261001T055349Z-ec722b14 $0.186 ($0.169–$0.203) vs $0.227: CHEAPER -18.1%, noise bound 2.8%, output length +9.1%
## Throttle check: Qwen/Qwen2.5-7B-Instruct

| | |
|---|---|
| Throttle version | 0.5.1 |
| Engine | vllm-rocm (from --config) |
| GPU | MI300X |
| Model | Qwen/Qwen2.5-7B-Instruct |
| GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) |
| Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) |
| Prompt cache | cold (unique prompt per request) |
| Label | fp8 |
| Check | 20261001T055349Z-ec722b14 (2026-10-01) |

**Result:** $0.1860 per million output tokens (95% CI $0.1691 to $0.2030, 5 blocks) [MEASURED]

**Compared with** check 20261001T054905Z-0f15c02b:

- before: $0.2272/M output tokens (95% CI $0.2217 to $0.2327)
- after: $0.1860/M output tokens (95% CI $0.1691 to $0.2030)
- change: -18.1%
- noise floor: calibrated: bound 2.8% (t x sqrt(2) x 0.6% run-to-run SD, 3 df)
- **Verdict: CHEAPER**, the -18.1% change is larger than the run-to-run noise bound (2.8% = t x sqrt(2) x 0.6% run-to-run SD, 3 df) and the 95% confidence intervals do not overlap

**What changed in config:**

- `config.quantization`: (not set) -> fp8
- `label`: baseline -> fp8
- `vllm.cache_config.num_gpu_blocks`: 186542 -> 193542

_Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._
Qwen2.5-7B-InstructvLLM 0.23.1 (ROCm) 1x AMD MI300X (Hot Aisle) vLLM defaults → --max-num-seqs 8 (at 32 concurrent requests) $0.227$0.222–$0.233 $0.700$0.684–$0.716 +208.0 to +213.4% MORE EXPENSIVE4 of 4

A deliberately bad setting for this load: shows what a harmless-looking flag costs.

Answers
Not sampled; output length within 2.5% of baseline.
GPU price
$2.99/hr ASSUMED Hot Aisle on-demand list price per MI300X GPU. Tokens and time are MEASURED.
Workload
5 blocks × 32 requests, concurrency 32, max 256 output tokens, temperature 0, cache: cold
Baseline repeats
$0.227, $0.230, $0.228, $0.227 (noise bound 2.8%)
Method
throttle check (throttle-pro 0.5.1), re-judged with the OUTPUT CHANGED rule. Date: 2026-10-01.
Source
the saved check history (summary below is Throttle's own sanitized --share-id output)

Every check of this change

  • 20261001T055113Z-acda5e76 $0.700 ($0.682–$0.718) vs $0.227: MORE EXPENSIVE +208.2%, noise bound 2.8%, output length -0.3%
  • 20261001T055624Z-9f099538 $0.701 ($0.681–$0.720) vs $0.227: MORE EXPENSIVE +208.4%, noise bound 2.8%, output length +0.4%
  • 20261001T055652Z-4340d2de $0.712 ($0.692–$0.732) vs $0.227: MORE EXPENSIVE +213.4%, noise bound 2.1%, output length -1.5%
  • 20261001T055720Z-7973bb09 $0.700 ($0.684–$0.716) vs $0.227: MORE EXPENSIVE +208.0%, noise bound 2.8%, output length -2.3%
## Throttle check: Qwen/Qwen2.5-7B-Instruct

| | |
|---|---|
| Throttle version | 0.5.1 |
| Engine | vllm-rocm (from --config) |
| GPU | MI300X |
| Model | Qwen/Qwen2.5-7B-Instruct |
| GPU hourly rate | $2.99/hr (ASSUMED, user-supplied) |
| Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) |
| Prompt cache | cold (unique prompt per request) |
| Label | seqs8 |
| Check | 20261001T055720Z-7973bb09 (2026-10-01) |

**Result:** $0.6998 per million output tokens (95% CI $0.6839 to $0.7158, 5 blocks) [MEASURED]

**Compared with** check 20261001T054905Z-0f15c02b:

- before: $0.2272/M output tokens (95% CI $0.2217 to $0.2327)
- after: $0.6998/M output tokens (95% CI $0.6839 to $0.7158)
- change: +208.0%
- noise floor: calibrated: bound 2.8% (t x sqrt(2) x 0.8% run-to-run SD, 5 df)
- **Verdict: MORE EXPENSIVE**, the +208.0% change is larger than the run-to-run noise bound (2.8% = t x sqrt(2) x 0.8% run-to-run SD, 5 df) and the 95% confidence intervals do not overlap

**What changed in config:**

- `config.max_num_seqs`: default -> 8
- `label`: baseline -> seqs8
- `vllm.cache_config.num_gpu_blocks`: 186542 -> 187972

_Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._
Qwen2.5-72B-InstructvLLM 0.23.1 (ROCm) 1x vs 2x AMD MI300X (Hot Aisle) 1 GPU (TP1, $2.99/hr) → 2 GPUs (--tensor-parallel-size 2, $5.98/hr), same BF16 weights $1.66$1.66–$1.67 $2.13$2.13–$2.14 +28.1 to +28.8% MORE EXPENSIVE3 of 3

Throttle 0.5.1 prints CHEAPER here because its verdict rescales to the baseline's GPU rate; this row judges real dollars with the same rule (2 GPUs: 1.56x the speed, 2x the price).

Answers
Output per block identical (6,320–6,333 tokens); 6-prompt sample identical on 5 of 6, sixth differs in wording only.
GPU price
$5.98/hr ASSUMED Hot Aisle on-demand list price $2.99 per MI300X GPU x GPUs used. Tokens and time are MEASURED.
Workload
5 blocks × 32 requests, concurrency 32, max 256 output tokens, temperature 0, cache: cold
Baseline repeats
$1.67, $1.67, $1.66, $1.66 (noise bound 0.8%)
Method
throttle check (throttle-pro 0.5.1), re-judged with the OUTPUT CHANGED rule. Date: 2026-10-01.
Source
the saved check history (summary below is Throttle's own sanitized --share-id output)

Every check of this change

  • 20261001T192537Z-5e84bc2d $2.14 ($2.13–$2.15) vs $1.66: MORE EXPENSIVE +28.8%, noise bound 1.0%, output length -0.1%
  • 20261001T192618Z-c00f5d81 $2.14 ($2.13–$2.15) vs $1.66: MORE EXPENSIVE +28.5%, noise bound 1.0%, output length -0.2%
  • 20261001T192700Z-d565005e $2.13 ($2.13–$2.14) vs $1.66: MORE EXPENSIVE +28.1%, noise bound 0.8%, output length -0.0%
## Throttle check: Qwen/Qwen2.5-72B-Instruct

| | |
|---|---|
| Throttle version | 0.5.1 |
| Engine | vllm (from --config) |
| GPU | MI300X |
| Model | Qwen/Qwen2.5-72B-Instruct |
| GPU hourly rate | $5.98/hr (ASSUMED, user-supplied) |
| Workload | 5 blocks x 32 requests, concurrency 32, max 256 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) |
| Prompt cache | cold (unique prompt per request) |
| Label | tp2 |
| Check | 20261001T192700Z-d565005e (2026-10-01) |

**Result:** $2.13 per million output tokens (95% CI $2.13 to $2.14, 5 blocks) [MEASURED]

**Compared with** check 20261001T192145Z-38a09569:

- before: $1.66/M output tokens (95% CI $1.66 to $1.67)
- after: $2.13/M output tokens (95% CI $2.13 to $2.14)
- change: +28.1% at the GPU rates as typed
- measured change at the baseline's GPU rate: +28.1%
- noise floor: calibrated: bound 0.8% (t x sqrt(2) x 0.2% run-to-run SD, 4 df)
- **Verdict: MORE EXPENSIVE**, in real dollars (GPU count and price changed) the +28.1% change is larger than the run-to-run noise bound (0.8%) and the 95% confidence intervals do not overlap

**What changed in config:**

- `config.gpus`: 1 -> 2
- `config.tensor_parallel`: 1 -> 2
- `gpu_hourly_rate_usd (ASSUMED)`: 2.9900 -> 5.9800
- `label`: tp1 -> tp2
- `vllm.cache_config.num_gpu_blocks`: 7776 -> 43167

_Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._
llama3.2:1bOllama 0.32.5 Apple M3 Pro (MacBook) OLLAMA_NUM_PARALLEL 1 → 4 $5.93$5.80–$6.06 $2.01$1.96–$2.06 -66.1 to -61.7% CHEAPER2 of 2

Laptop; the $/hr is an assumed rate for illustration.

Answers
Not sampled.
GPU price
$1.50/hr ASSUMED assumed $1.50/hr (laptop, no real bill). Tokens and time are MEASURED.
Workload
5 blocks × 8 requests, concurrency 4, max 128 output tokens, temperature 0, cache: cold
Baseline repeats
$6.42, $6.22, $6.10, $5.93 (noise bound 15.0%)
Method
throttle check (throttle-pro 0.5.1), re-judged with the OUTPUT CHANGED rule. Date: 2026-09-30.
Source
the saved check history (summary below is Throttle's own sanitized --share-id output)

Every check of this change

  • 20260930T064318Z-aa845aa6 $2.27 ($1.99–$2.54) vs $5.93: CHEAPER -61.7%, noise bound 15.0%, output length -2.1%
  • 20260930T064345Z-bef600a8 $2.01 ($1.96–$2.06) vs $5.93: CHEAPER -66.1%, noise bound 15.0%, output length +0.0%
## Throttle check: llama3.2:1b

| | |
|---|---|
| Throttle version | 0.5.1 |
| Engine | Ollama (likely; /v1/models owned_by=library) |
| GPU | not given (add --config gpu=NAME) |
| Model | llama3.2:1b |
| GPU hourly rate | $1.50/hr (ASSUMED, user-supplied) |
| Workload | 5 blocks x 8 requests, concurrency 4, max 128 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) |
| Prompt cache | cold (unique prompt per request) |
| Label | parallel-4 |
| Check | 20260930T064345Z-bef600a8 (2026-09-30) |

**Result:** $2.01 per million output tokens (95% CI $1.96 to $2.06, 5 blocks) [MEASURED]

**Compared with** check 20260930T064223Z-93b28975:

- before: $5.93/M output tokens (95% CI $5.80 to $6.06)
- after: $2.01/M output tokens (95% CI $1.96 to $2.06)
- change: -66.1%
- noise floor: calibrated: bound 15.0% (t x sqrt(2) x 3.3% run-to-run SD, 3 df)
- **Verdict: CHEAPER**, the -66.1% change is larger than the run-to-run noise bound (15.0% = t x sqrt(2) x 3.3% run-to-run SD, 3 df) and the 95% confidence intervals do not overlap

**What changed in config:**

- `config.ollama_num_parallel`: 1 -> 4
- `label`: parallel-1 -> parallel-4

_Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._
llama3.2:3bOllama 0.32.5 Apple M3 Pro (MacBook) Q4_K_M → Q8_0 weights (OLLAMA_NUM_PARALLEL=4) $5.18$5.08–$5.27 $4.70$4.53–$4.87 -10.3 to -9.2% CHEAPER3 of 3

Laptop; the $/hr is an assumed rate. Q4/Q8 order was interleaved to rule out drift.

Answers
Not sampled.
GPU price
$1.50/hr ASSUMED assumed $1.50/hr (laptop, no real bill). Tokens and time are MEASURED.
Workload
5 blocks × 8 requests, concurrency 4, max 128 output tokens, temperature 0, cache: cold
Baseline repeats
$5.17, $5.22, $5.28, $5.18 (noise bound 3.9%)
Method
throttle check (throttle-pro 0.5.1), re-judged with the OUTPUT CHANGED rule. Date: 2026-10-01.
Source
the saved check history (summary below is Throttle's own sanitized --share-id output)

Every check of this change

  • 20261001T034603Z-393086ba $4.64 ($4.54–$4.74) vs $5.18: CHEAPER -10.3%, noise bound 4.5%, output length +0.0%
  • 20261001T034701Z-3909ebcc $4.69 ($4.42–$4.95) vs $5.18: CHEAPER -9.4%, noise bound 4.5%, output length +0.3%
  • 20261001T035022Z-29452cbf $4.70 ($4.53–$4.87) vs $5.18: CHEAPER -9.2%, noise bound 3.9%, output length -0.3%
## Throttle check: llama3.2:3b-instruct-q8_0

| | |
|---|---|
| Throttle version | 0.5.1 |
| Engine | Ollama (likely; /v1/models owned_by=library) |
| GPU | not given (add --config gpu=NAME) |
| Model | llama3.2:3b-instruct-q8_0 |
| GPU hourly rate | $1.50/hr (ASSUMED, user-supplied) |
| Workload | 5 blocks x 8 requests, concurrency 4, max 128 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) |
| Prompt cache | cold (unique prompt per request) |
| Label | q8 |
| Check | 20261001T035022Z-29452cbf (2026-10-01) |

**Result:** $4.70 per million output tokens (95% CI $4.53 to $4.87, 5 blocks) [MEASURED]

**Compared with** check 20261001T034456Z-06d7816a:

- before: $5.18/M output tokens (95% CI $5.08 to $5.27)
- after: $4.70/M output tokens (95% CI $4.53 to $4.87)
- change: -9.2%
- noise floor: calibrated: bound 3.9% (t x sqrt(2) x 1.1% run-to-run SD, 6 df)
- **Verdict: CHEAPER**, the -9.2% change is larger than the run-to-run noise bound (3.9% = t x sqrt(2) x 1.1% run-to-run SD, 6 df) and the 95% confidence intervals do not overlap

**What changed in config:**

- `config.quant`: Q4_K_M -> Q8_0
- `label`: q4 -> q8
- `request.model`: llama3.2:3b -> llama3.2:3b-instruct-q8_0

_Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._
llama3.2:3b → llama3.2:1bOllama Apple M3 Pro (MacBook) Swap llama3.2:3b for llama3.2:1b $8.63$8.27–$9.00 $5.91$4.85–$6.97 -35.9 to -31.6% CHEAPER2 of 2

Laptop; the $/hr is an assumed rate.

Answers
Not sampled. A smaller model is a quality trade-off Throttle does not judge.
GPU price
$1.50/hr ASSUMED assumed $1.50/hr (laptop, no real bill). Tokens and time are MEASURED.
Workload
3 blocks × 4 requests, concurrency 2, max 64 output tokens, temperature 0, cache: cold
Baseline repeats
$8.49, $8.68, $8.71, $8.63 (noise bound 5.0%)
Method
throttle check (throttle-pro 0.4.2), re-judged with the OUTPUT CHANGED rule. Date: 2026-09-27.
Source
the saved check history (summary below is Throttle's own sanitized --share-id output)

Every check of this change

  • 20260927T194325Z-40935488 $5.54 ($5.51–$5.57) vs $8.63: CHEAPER -35.9%, noise bound 5.0%, output length +0.0%
  • 20260927T194337Z-41f5e1f0 $5.91 ($4.85–$6.97) vs $8.63: CHEAPER -31.6%, noise bound 5.0%, output length +0.0%
## Throttle check: llama3.2:1b

| | |
|---|---|
| Throttle version | 0.4.2 |
| Engine | Ollama (likely; /v1/models owned_by=library) |
| GPU | apple-m3-pro |
| Model | llama3.2:1b |
| GPU hourly rate | $1.50/hr (ASSUMED, user-supplied) |
| Workload | 3 blocks x 4 requests, concurrency 2, max 64 output tokens, temperature 0, 8 built-in prompts (sha256 787fc93d62e1) |
| Prompt cache | cold (unique prompt per request) |
| Label | smaller-1b |
| Check | 20260927T194337Z-41f5e1f0 (2026-09-27) |

**Result:** $5.91 per million output tokens (95% CI $4.85 to $6.97, 3 blocks) [MEASURED]

**Compared with** check 20260927T194311Z-48d7dd1d:

- before: $8.63/M output tokens (95% CI $8.27 to $9.00)
- after: $5.91/M output tokens (95% CI $4.85 to $6.97)
- change: -31.6%
- noise floor: calibrated: bound 5.0% (t x sqrt(2) x 1.1% run-to-run SD, 3 df)
- **Verdict: CHEAPER**, the -31.6% change is larger than the run-to-run noise bound (5.0% = t x sqrt(2) x 1.1% run-to-run SD, 3 df) and the 95% confidence intervals do not overlap

**What changed in config:**

- `label`: baseline-3b -> smaller-1b
- `request.model`: llama3.2:3b -> llama3.2:1b

_Shared with `throttle check --share`. No URLs, hostnames, IPs or keys are included; config values that look like secrets are masked._
Qwen2.5-0.5B-InstructvLLM 0.16.0 1x NVIDIA A100 80GB PCIe (RunPod) max_num_seqs 1 → 8 $0.746 $0.234 -68.6% CHEAPERdecision-eligible (golden protocol)

A deliberately bad baseline (one sequence at a time) on a small model; it shows the measurement, not a saving to expect.

Answers
Not sampled.
GPU price
$1.39/hr ASSUMED RunPod on-demand rate at run time. Tokens and time are MEASURED.
Workload
3 blocks × 67 requests, concurrency 8, max 128 output tokens, temperature 0, cache: prefix caching disabled
Baseline repeats
$0.748, $0.746, $0.745
Method
throttle golden: six-position counterbalanced protocol (3 baseline, 3 candidate positions), decision_eligible: true. Date: 2026-08-17.

$ per million output tokens = measured wall-clock time × the GPU's hourly price ÷ measured output tokens. Ranges show every check of a change; CIs are 95%. One workload per row: your prompts, concurrency and output lengths will give different numbers. Download the data (JSON).

Add your result

The Index grows from real checks. If you serve a model on your own or rented GPUs, measure one change and send it in:

  1. Install Throttle: pipx install throttle-pro
  2. Run throttle check 3 times on your current config to measure your noise, then once after your change.
  3. Read a few answers from both configs (the FP8 rows above are why).
  4. Run throttle check --share-id <id> and paste the summary into a new Index submission. The summary leaves out URLs, hostnames and anything that looks like a secret.

Every accepted row names its GPU, engine, workload and the $/hr it assumed. Nothing is added without the check output behind it.

You pick, we measure

Every week the community votes on one serving question, and we rent the GPU, run it with Throttle and add the result here, including when the answer is "no winner." Want a question measured? Open a request on GitHub or email kushthrottle@gmail.com.