LLM cost per million tokens: the formula, a calculator, and real numbers
If you serve your own model, every token has a price. Here is the formula, a calculator you can use now, and a worked example from a real run.
The formula
A self-hosted GPU costs the same per hour whether it is busy or idle. What changes with your config is how many tokens it produces in that hour. So cost per million output tokens is:
$ per million tokens = GPU $ per hour ÷ (tokens per second × 3600) × 1,000,000
Two inputs. The hourly price is what you pay for every GPU the model runs on. Tokens per second is the aggregate output throughput of the server at your real concurrency, not one request's streaming speed.
Calculator
The defaults are a real measurement: Qwen2.5-7B-Instruct on one AMD MI300X with vLLM, 32 concurrent requests, at Hot Aisle's $2.99/hr list price. The tokens per second value is the hard part: it has to be measured on your stack, at your concurrency.
Worked example: 7B and 72B on an MI300X
| Setup | $/hr | Output tok/s | $ per M output tokens |
|---|---|---|---|
| Qwen2.5-7B, 1× MI300X | $2.99 | ~3,656 | $0.227 |
| Qwen2.5-72B, 1× MI300X | $2.99 | ~498 | $1.67 |
| Qwen2.5-72B, 2× MI300X (tensor parallel 2) | $5.98 | ~777 | $2.14 |
Check the first row: 2.99 ÷ (3,656 × 3,600) × 1,000,000 = $0.227. The last row shows why the formula matters: two GPUs made the 72B model 1.56× faster but each token 28% more expensive, because the hourly price doubled. More on that.
Four ways the number goes wrong
1. Measuring one request. A single stream at 40 tok/s says nothing about a server batching 32 requests. Measure at the concurrency you run in production.
2. A warm cache. Send the same prompt repeatedly and a prefix cache serves most of it, so your cost looks lower than real traffic will be. Throttle tags each request with a unique prefix by default (cold cache) so a cache can't flatter the result.
3. Ignoring noise. The same untouched server can move a lot between runs. Four checks of one unchanged Ollama server came back at $8.77, $8.86, $11.26 and $10.02 per million output tokens. A 10% "saving" inside that spread is not a saving. The full story.
4. Forgetting the answers. A cheaper token is not a saving if the output changed. On one run, vLLM's on-the-fly FP8 made a 32B model look 41% cheaper while it answered with "!!!!!!". See the FP8 guide.
Measure it instead of estimating it
The calculator is only as good as the throughput you type in. throttle check measures it against your own OpenAI-compatible endpoint and prints $ per million tokens with a 95% confidence interval:
pipx install throttle-pro
throttle check \
--url http://localhost:8000 \
--model Qwen/Qwen2.5-7B-Instruct \
--gpu-hourly-rate 2.99
Run it again after any config change and it compares the two, calling a winner only when the change beats your measured run-to-run noise. For the step-by-step on vLLM, SGLang and Ollama, see how to measure cost per million tokens on a self-hosted LLM.
FAQ
How do you calculate LLM cost per million tokens?
Divide the hourly price of every GPU the model uses by the server's aggregate output tokens per hour, then multiply by one million: $/hr ÷ (tokens/s × 3600) × 1,000,000.
What is a typical cost per million tokens for a self-hosted model?
It depends on the model, GPU and concurrency. Measured examples on one AMD MI300X at $2.99/hr: Qwen2.5-7B about $0.227 and Qwen2.5-72B about $1.67 per million output tokens, at 32 concurrent requests.
Should I count input tokens too?
Throttle reports $/M for output tokens, input tokens and total tokens from the same GPU spend. Output tokens are usually what you compare, since decoding dominates GPU time for chat-style workloads.
Why does my cost per token change between runs?
Load, thermals, other tenants and batching vary. Measure the run-to-run noise with repeat checks before trusting a small difference.
Related guides
vLLM cost per token: measure it, then change one flag at a time
Measure vLLM cost per million tokens on your own GPU, then test max-num-seqs, tensor parallel size and FP8 one at a time. Real results from AMD MI300X and NVIDIA A100 runs.
GuideSelf-hosted LLM cost: how to compare it with an API, honestly
What does a self-hosted LLM really cost? How to turn GPU rental prices into dollars per million tokens, why utilization decides the answer, and real measurements on AMD MI300X.