New: the Throttle Cost Index →

Cut your LLMcost per token.

Qwen2.5-72B on one MI300X: $1.67 per million tokens. On two: $2.14. Measured, not guessed.

See the Cost Index →
Watch the 75-second film

Every config change is a new bill. Throttle checks it.

  1. Measure

    throttle check prints your $ per million tokens, with a 95% interval.

  2. Change one thing

    A flag, the model, the engine, the quantization.

  3. Get a verdict

    Only when the change beats your measured run-to-run noise.

    CHEAPERMORE EXPENSIVENO WINNER

Real runs. Real dollars.

GPU count · AMD MI300X

One GPU beat two.

$1.67vs$2.14
per M output tokens · 1 GPU vs 2

Two GPUs ran 1.56× faster but cost 28% more per token.

Qwen2.5-72B BF16, vLLM 0.23 ROCm, 32 concurrent requests. 4 + 3 checks, Oct 1 2026.

one vLLM flag · AMD MI300X

One flag tripled the bill.

$0.227→$0.700
per M output tokens · +208%

--max-num-seqs 8 at 32 concurrent requests.

Qwen2.5-7B, vLLM 0.23 ROCm. MORE EXPENSIVE in 4 of 4 checks.

one vLLM flag · NVIDIA A100

One flag cut it by two-thirds.

$0.746→$0.234
per M output tokens · −68.6%

max_num_seqs 1 → 8.

Qwen2.5-0.5B, vLLM 0.16, A100 80GB. Counterbalanced golden run.

GPU $/hr are list prices (assumed). Tokens and time are measured.Every measured result → Cost Index

Real terminal runs of throttle-pro 0.4.2 on a MacBook (GPU rate assumed at $1.50/hr), and the A100 result from the saved golden run.

Want your number? We’ll measure it with you.

cost audit · 2 weeks

Your true $ per million tokens

$500one-time

Measured on your stack, plus verdicts on up to 3 changes and a CI gate that fails a costlier deploy.

Book a 15-minute call →

No savings guarantee. What the audit covers

pro · early access

Scheduled checks and alerts

$19/ month, launch price