All posts Guide

How to measure cost per million tokens on a self-hosted LLM (vLLM, SGLang, Ollama)

A practical guide to turning GPU dollars per hour into dollars per million tokens for vLLM, SGLang and Ollama, the four pitfalls that make the number wrong, and how to re-check it after every config change and in CI.

API providers price per million tokens. When you self-host, you pay per GPU hour. To compare the two, or to know whether last week's config change helped, you need your own number in the provider's unit: dollars per million tokens, measured on your server, at your load.

This guide covers the math, the four mistakes that make the number wrong, and a step-by-step way to measure it on vLLM, SGLang or Ollama. The commands use Throttle, an open-source CLI I build, but the method works with any load generator and a stopwatch.

The formula

$/M tokens = (GPU $/hour × number of GPUs × wall-clock hours) ÷ tokens × 1,000,000

Equivalently, from throughput:

$/M output tokens = GPU $/hour ÷ (3,600 × output tokens per second) × 1,000,000

A real example: one A100 80GB PCIe rented at $1.39/hour, serving Qwen2.5-0.5B-Instruct on vLLM 0.16.0 at 516.55 output tokens per second:

1.39 ÷ (3,600 × 516.55) × 1,000,000 = $0.7475 per million output tokens

With one flag changed, the same GPU produced 1,632.41 tokens per second, and the same formula gives $0.2371. That's the whole game: the hourly price is fixed, so cost per token is throughput in disguise, and anything that moves throughput moves your bill.

Which tokens?

  • Output $/M charges the full GPU time to output tokens. It's the usual headline because decode dominates most chat workloads.
  • Input $/M charges the same time to input tokens.
  • Blended $/M divides by input plus output tokens. This is the one to hold next to a provider's per-token bill.

Don't add input $/M and output $/M together. Each one already charges the whole run to one token type, so adding them counts the GPU time twice.

Four pitfalls that make the number wrong

1. Measuring at the wrong load. Throughput, and so $/M, depends on how many requests the server works on at once. One request at a time on a big GPU leaves most of it idle and makes every token look expensive. Measure at the concurrency you actually run in production, and state it next to the number.

2. A warm prefix cache. vLLM (--enable-prefix-caching, on by default in recent versions) and SGLang (RadixAttention) reuse work for prompts they've seen before. If your benchmark sends the same prompts every run, later runs skip prefill and look cheaper than your real traffic will be. Either vary the prompts or tag each request. Since 0.4.1, Throttle prefixes every request with a unique [run NNNNNN req NNNN] tag so the cache stays cold, and --warm-cache measures a warm cache on purpose.

3. Believing one run. Two runs of the same unchanged server can disagree by a lot. In one session on my laptop, three unchanged checks spread 9.8%. In another, one unchanged run came in 27% higher than the one before it (the full story). A before/after pair can't separate a real change from load. You have to measure run-to-run noise first.

4. Treating the GPU price as a measurement. The hourly rate is an input you choose: on-demand, spot, reserved, or amortized hardware you own. Label it as an assumption. Tokens and seconds are the only things you actually measure.

Step by step

0. Install, and try it with no GPU

pipx install throttle-pro
throttle demo      # no GPU needed; every number is SIMULATED

1. Point it at your server

Any OpenAI-compatible endpoint works. Default ports:

Engine URL
vLLM http://localhost:8000
SGLang http://localhost:30000
Ollama http://localhost:11434

If the endpoint needs a key, put it in an environment variable and pass the variable's name with --api-key-env, never the key itself.

2. Get a first number

throttle cost --url http://localhost:11434 --model llama3.2:3b \
  --gpu-hourly-rate 1.50 --num-requests 5

Recorded on a MacBook with local Ollama and llama3.2:3b. The GPU rate is assumed, since a laptop has no GPU invoice:

Measured Cost [MEASURED time x ASSUMED $1.50/hr]:
  GPU hours: 0.001494
  Total cost: $0.0022
  Blended cost: $2.87 per million tokens (input + output, 780 tokens)
  Input cost: $3.78 per million tokens
  Output cost: $11.98 per million tokens

This is a quick spot price, not a decision-grade number: five sequential requests, one synthetic prompt.

3. Measure it properly, three times, before you change anything

throttle check measures in repeated blocks, reports a 95% confidence interval, fingerprints the serving config it can see, and saves every result so the next check is compared with it.

for i in 1 2 3; do
  throttle check --url http://localhost:11434 --model llama3.2:3b \
    --gpu-hourly-rate 1.50 --label before
done

The first check has nothing to compare with:

Result
  $8.64 per million output tokens   [MEASURED]
  95% CI $8.08 to $9.20, across 5 blocks (Student t)

Checks two and three come back NOT CALIBRATED on purpose: until there are three earlier checks of one config, there's no estimate of run-to-run noise to judge a change against.

On vLLM, Throttle also reads config labels from /metrics. For anything the endpoint doesn't report, such as the GPU type or a flag you set at launch, declare it with --config:

--config gpu=A100-80GB --config max_num_seqs=8

4. Change one thing, then check again

throttle check --url http://localhost:11434 --model llama3.2:3b \
  --gpu-hourly-rate 1.50 --label after --config OLLAMA_NUM_PARALLEL=4

In this recording the server was not actually restarted, so nothing really changed, and the check still measured 11.5% cheaper. The verdict:

  change  -$1.09 per million output tokens (-11.5%)
  noise   run-to-run noise bound 30.4% [MEASURED from 3 earlier checks of
          unchanged configs within 24 h: t x sqrt(2) x 5.0% SD, 2 df]
  monthly not projected: with NO WINNER the difference is not distinguishable from noise
  Verdict: NO WINNER, the change (-11.5%) is not larger than the run-to-run
  noise bound (30.4% = t x sqrt(2) x 5.0% run-to-run SD, 2 df); and the 95%
  confidence intervals overlap, so the difference is within measurement noise.

The rules: CHEAPER or MORE EXPENSIVE only when the change is larger than the run-to-run noise bound and the 95% intervals don't overlap. Otherwise it's NO WINNER. With --monthly-tokens, a calibrated verdict also prints the monthly dollar difference, labelled PROJECTED.

5. Make it a gate in CI

Run the check after every deploy. --fail-if-costlier PCT turns the verdict into an exit code:

Exit code Meaning
0 Check done: CHEAPER, NO WINNER, or the first check
1 Measurement failed; nothing recorded
4 A calibrated MORE EXPENSIVE by at least PCT percent
5 A threshold is set but the change couldn't be judged (NOT CALIBRATED, or the baseline used a different workload)

The repo ships a GitHub Action that runs the check, caches the check history between runs, and writes the verdict to the job summary:

- name: Throttle cost check
  uses: KushagraKanaujia/throttle@main   # pin a release tag or commit SHA in production
  env:
    VLLM_API_KEY: ${{ secrets.VLLM_API_KEY }}
  with:
    url: https://vllm.internal.example:8000
    model: meta-llama/Llama-3.1-8B-Instruct
    gpu-hourly-rate: "2.49"
    api-key-env: VLLM_API_KEY
    config: |
      gpu=H100
      max_num_seqs=256
    concurrency: "8"            # match production load
    fail-if-costlier: "5"       # fail on a calibrated +5% or worse
    not-calibrated: warn

The runner has to be able to reach the endpoint, which usually means a self-hosted runner on the same network. The first three runs are NOT CALIBRATED by design; after that, a deploy that makes every token measurably more expensive fails the job instead of showing up on next month's bill. The full setup is in docs/github-action.md.

Checklist

  • Hourly rate written down and labelled as an assumption
  • Measured at production concurrency
  • Prefix cache cold (or warm on purpose, and said so)
  • Unchanged config measured at least 3 times before the change
  • Change believed only if it beats the noise bound and the intervals don't overlap
  • Re-checked automatically after every deploy

Throttle is MIT-licensed and runs entirely on your machine: github.com/KushagraKanaujia/throttle. If you measure something interesting, throttle check --share prints a sanitized summary you can post as a results issue.

Keep reading