How to measure cost per million tokens on a self-hosted LLM (vLLM, SGLang, Ollama)
A practical guide to turning GPU dollars per hour into dollars per million tokens for vLLM, SGLang and Ollama, the four pitfalls that make the number wrong, and how to re-check it after every config change and in CI.
API providers price per million tokens. When you self-host, you pay per GPU hour. To compare the two, or to know whether last week's config change helped, you need your own number in the provider's unit: dollars per million tokens, measured on your server, at your load.
This guide covers the math, the four mistakes that make the number wrong, and a step-by-step way to measure it on vLLM, SGLang or Ollama. The commands use Throttle, an open-source CLI I build, but the method works with any load generator and a stopwatch.
The formula
$/M tokens = (GPU $/hour × number of GPUs × wall-clock hours) ÷ tokens × 1,000,000
Equivalently, from throughput:
$/M output tokens = GPU $/hour ÷ (3,600 × output tokens per second) × 1,000,000
A real example: one A100 80GB PCIe rented at $1.39/hour, serving Qwen2.5-0.5B-Instruct on vLLM 0.16.0 at 516.55 output tokens per second:
1.39 ÷ (3,600 × 516.55) × 1,000,000 = $0.7475 per million output tokens
With one flag changed, the same GPU produced 1,632.41 tokens per second, and the same formula gives $0.2371. That's the whole game: the hourly price is fixed, so cost per token is throughput in disguise, and anything that moves throughput moves your bill.
Which tokens?
- Output $/M charges the full GPU time to output tokens. It's the usual headline because decode dominates most chat workloads.
- Input $/M charges the same time to input tokens.
- Blended $/M divides by input plus output tokens. This is the one to hold next to a provider's per-token bill.
Don't add input $/M and output $/M together. Each one already charges the whole run to one token type, so adding them counts the GPU time twice.
Four pitfalls that make the number wrong
1. Measuring at the wrong load. Throughput, and so $/M, depends on how many requests the server works on at once. One request at a time on a big GPU leaves most of it idle and makes every token look expensive. Measure at the concurrency you actually run in production, and state it next to the number.
2. A warm prefix cache. vLLM (--enable-prefix-caching, on by default in recent versions) and SGLang (RadixAttention) reuse work for prompts they've seen before. If your benchmark sends the same prompts every run, later runs skip prefill and look cheaper than your real traffic will be. Either vary the prompts or tag each request. Since 0.4.1, Throttle prefixes every request with a unique [run NNNNNN req NNNN] tag so the cache stays cold, and --warm-cache measures a warm cache on purpose.
3. Believing one run. Two runs of the same unchanged server can disagree by a lot. In one session on my laptop, three unchanged checks spread 9.8%. In another, one unchanged run came in 27% higher than the one before it (the full story). A before/after pair can't separate a real change from load. You have to measure run-to-run noise first.
4. Treating the GPU price as a measurement. The hourly rate is an input you choose: on-demand, spot, reserved, or amortized hardware you own. Label it as an assumption. Tokens and seconds are the only things you actually measure.
Step by step
0. Install, and try it with no GPU
pipx install throttle-pro
throttle demo # no GPU needed; every number is SIMULATED
1. Point it at your server
Any OpenAI-compatible endpoint works. Default ports:
| Engine | URL |
|---|---|
| vLLM | http://localhost:8000 |
| SGLang | http://localhost:30000 |
| Ollama | http://localhost:11434 |
If the endpoint needs a key, put it in an environment variable and pass the variable's name with --api-key-env, never the key itself.
2. Get a first number
throttle cost --url http://localhost:11434 --model llama3.2:3b \
--gpu-hourly-rate 1.50 --num-requests 5
Recorded on a MacBook with local Ollama and llama3.2:3b. The GPU rate is assumed, since a laptop has no GPU invoice:
Measured Cost [MEASURED time x ASSUMED $1.50/hr]:
GPU hours: 0.001494
Total cost: $0.0022
Blended cost: $2.87 per million tokens (input + output, 780 tokens)
Input cost: $3.78 per million tokens
Output cost: $11.98 per million tokens
This is a quick spot price, not a decision-grade number: five sequential requests, one synthetic prompt.
3. Measure it properly, three times, before you change anything
throttle check measures in repeated blocks, reports a 95% confidence interval, fingerprints the serving config it can see, and saves every result so the next check is compared with it.
for i in 1 2 3; do
throttle check --url http://localhost:11434 --model llama3.2:3b \
--gpu-hourly-rate 1.50 --label before
done
The first check has nothing to compare with:
Result
$8.64 per million output tokens [MEASURED]
95% CI $8.08 to $9.20, across 5 blocks (Student t)
Checks two and three come back NOT CALIBRATED on purpose: until there are three earlier checks of one config, there's no estimate of run-to-run noise to judge a change against.
On vLLM, Throttle also reads config labels from /metrics. For anything the endpoint doesn't report, such as the GPU type or a flag you set at launch, declare it with --config:
--config gpu=A100-80GB --config max_num_seqs=8
4. Change one thing, then check again
throttle check --url http://localhost:11434 --model llama3.2:3b \
--gpu-hourly-rate 1.50 --label after --config OLLAMA_NUM_PARALLEL=4
In this recording the server was not actually restarted, so nothing really changed, and the check still measured 11.5% cheaper. The verdict:
change -$1.09 per million output tokens (-11.5%)
noise run-to-run noise bound 30.4% [MEASURED from 3 earlier checks of
unchanged configs within 24 h: t x sqrt(2) x 5.0% SD, 2 df]
monthly not projected: with NO WINNER the difference is not distinguishable from noise
Verdict: NO WINNER, the change (-11.5%) is not larger than the run-to-run
noise bound (30.4% = t x sqrt(2) x 5.0% run-to-run SD, 2 df); and the 95%
confidence intervals overlap, so the difference is within measurement noise.
The rules: CHEAPER or MORE EXPENSIVE only when the change is larger than the run-to-run noise bound and the 95% intervals don't overlap. Otherwise it's NO WINNER. With --monthly-tokens, a calibrated verdict also prints the monthly dollar difference, labelled PROJECTED.
5. Make it a gate in CI
Run the check after every deploy. --fail-if-costlier PCT turns the verdict into an exit code:
| Exit code | Meaning |
|---|---|
| 0 | Check done: CHEAPER, NO WINNER, or the first check |
| 1 | Measurement failed; nothing recorded |
| 4 | A calibrated MORE EXPENSIVE by at least PCT percent |
| 5 | A threshold is set but the change couldn't be judged (NOT CALIBRATED, or the baseline used a different workload) |
The repo ships a GitHub Action that runs the check, caches the check history between runs, and writes the verdict to the job summary:
- name: Throttle cost check
uses: KushagraKanaujia/throttle@main # pin a release tag or commit SHA in production
env:
VLLM_API_KEY: ${{ secrets.VLLM_API_KEY }}
with:
url: https://vllm.internal.example:8000
model: meta-llama/Llama-3.1-8B-Instruct
gpu-hourly-rate: "2.49"
api-key-env: VLLM_API_KEY
config: |
gpu=H100
max_num_seqs=256
concurrency: "8" # match production load
fail-if-costlier: "5" # fail on a calibrated +5% or worse
not-calibrated: warn
The runner has to be able to reach the endpoint, which usually means a self-hosted runner on the same network. The first three runs are NOT CALIBRATED by design; after that, a deploy that makes every token measurably more expensive fails the job instead of showing up on next month's bill. The full setup is in docs/github-action.md.
Checklist
- Hourly rate written down and labelled as an assumption
- Measured at production concurrency
- Prefix cache cold (or warm on purpose, and said so)
- Unchanged config measured at least 3 times before the change
- Change believed only if it beats the noise bound and the intervals don't overlap
- Re-checked automatically after every deploy
Throttle is MIT-licensed and runs entirely on your machine: github.com/KushagraKanaujia/throttle. If you measure something interesting, throttle check --share prints a sanitized summary you can post as a results issue.
Keep reading
I changed nothing and my LLM server got 27% more expensive
Four cost checks of the same untouched Ollama server came back at $8.77, $8.86, $11.26 and $10.02 per million output tokens. Why same-server runs swing, and the noise bound I now use before believing any before/after number.
Case studyOne vLLM flag, −68.6% cost per token: what a counterbalanced benchmark looks like
A real A100 run where changing vLLM's max_num_seqs from 1 to 8 took cost from $0.746 to $0.234 per million output tokens, measured with a six-position B‑C‑B / C‑B‑C protocol. How the protocol works, the raw numbers, and what the result does not mean.