All guides Calculator

LLM cost per million tokens: the formula, a calculator, and real numbers

If you serve your own model, every token has a price. Here is the formula, a calculator you can use now, and a worked example from a real run.

The formula

A self-hosted GPU costs the same per hour whether it is busy or idle. What changes with your config is how many tokens it produces in that hour. So cost per million output tokens is:

$ per million tokens = GPU $ per hour ÷ (tokens per second × 3600) × 1,000,000

Two inputs. The hourly price is what you pay for every GPU the model runs on. Tokens per second is the aggregate output throughput of the server at your real concurrency, not one request's streaming speed.

Calculator

The defaults are a real measurement: Qwen2.5-7B-Instruct on one AMD MI300X with vLLM, 32 concurrent requests, at Hot Aisle's $2.99/hr list price. The tokens per second value is the hard part: it has to be measured on your stack, at your concurrency.

Worked example: 7B and 72B on an MI300X

Setup$/hrOutput tok/s$ per M output tokens
Qwen2.5-7B, 1× MI300X$2.99~3,656$0.227
Qwen2.5-72B, 1× MI300X$2.99~498$1.67
Qwen2.5-72B, 2× MI300X (tensor parallel 2)$5.98~777$2.14

Check the first row: 2.99 ÷ (3,656 × 3,600) × 1,000,000 = $0.227. The last row shows why the formula matters: two GPUs made the 72B model 1.56× faster but each token 28% more expensive, because the hourly price doubled. More on that.

Four ways the number goes wrong

1. Measuring one request. A single stream at 40 tok/s says nothing about a server batching 32 requests. Measure at the concurrency you run in production.

2. A warm cache. Send the same prompt repeatedly and a prefix cache serves most of it, so your cost looks lower than real traffic will be. Throttle tags each request with a unique prefix by default (cold cache) so a cache can't flatter the result.

3. Ignoring noise. The same untouched server can move a lot between runs. Four checks of one unchanged Ollama server came back at $8.77, $8.86, $11.26 and $10.02 per million output tokens. A 10% "saving" inside that spread is not a saving. The full story.

4. Forgetting the answers. A cheaper token is not a saving if the output changed. On one run, vLLM's on-the-fly FP8 made a 32B model look 41% cheaper while it answered with "!!!!!!". See the FP8 guide.

Measure it instead of estimating it

The calculator is only as good as the throughput you type in. throttle check measures it against your own OpenAI-compatible endpoint and prints $ per million tokens with a 95% confidence interval:

pipx install throttle-pro

throttle check \
  --url http://localhost:8000 \
  --model Qwen/Qwen2.5-7B-Instruct \
  --gpu-hourly-rate 2.99

Run it again after any config change and it compares the two, calling a winner only when the change beats your measured run-to-run noise. For the step-by-step on vLLM, SGLang and Ollama, see how to measure cost per million tokens on a self-hosted LLM.

FAQ

How do you calculate LLM cost per million tokens?

Divide the hourly price of every GPU the model uses by the server's aggregate output tokens per hour, then multiply by one million: $/hr ÷ (tokens/s × 3600) × 1,000,000.

What is a typical cost per million tokens for a self-hosted model?

It depends on the model, GPU and concurrency. Measured examples on one AMD MI300X at $2.99/hr: Qwen2.5-7B about $0.227 and Qwen2.5-72B about $1.67 per million output tokens, at 32 concurrent requests.

Should I count input tokens too?

Throttle reports $/M for output tokens, input tokens and total tokens from the same GPU spend. Output tokens are usually what you compare, since decoding dominates GPU time for chat-style workloads.

Why does my cost per token change between runs?

Load, thermals, other tenants and batching vary. Measure the run-to-run noise with repeat checks before trusting a small difference.

Related guides