All guides Guide

Self-hosted LLM cost: how to compare it with an API, honestly

An API price is per token. A GPU price is per hour. To compare them you need one number your provider won't give you: your own cost per million tokens.

Two kinds of price

A hosted API charges per million tokens, whether you send one request a day or a thousand a second. A rented or owned GPU charges per hour, whether it is busy or idle. The only way to compare them is to convert the GPU's hourly price into a price per token, using the throughput your server actually achieves:

$ per million tokens = GPU $ per hour ÷ (tokens per second × 3600) × 1,000,000

Try it with your numbers in the calculator.

Utilization decides the answer

That formula assumes the GPU is serving traffic the whole hour. Real servers aren't. If your server is busy 30% of the time, your effective cost per token is about 3.3× the measured one, because you pay for the idle 70% too. This is the main reason teams who self-host for cost savings are sometimes surprised: the per-token cost under load can be very low, and the per-token cost averaged over the month much higher.

So you need two numbers: cost per token while serving (measured), and how much of the time you are serving (from your traffic). Multiply the first by 1 ÷ utilization.

Measured examples

ModelHardwareConfig$ per M output tokens
Qwen2.5-7B1× MI300X, $2.99/hrvLLM defaults$0.227
Qwen2.5-7B1× MI300XPre-quantized FP8$0.186
Qwen2.5-72B1× MI300XvLLM defaults, BF16$1.67
Qwen2.5-72B1× MI300XPre-quantized FP8$1.13
Qwen2.5-72B2× MI300X, $5.98/hrTensor parallel 2$2.14

All at 32 concurrent requests and 256 max output tokens, i.e. a busy server. These are the "while serving" numbers; adjust for your utilization. Compare them with the per-million price your API provider lists for a comparable model, and remember API prices include their idle capacity, reliability and operations.

Costs the formula leaves out

The GPU hour is the biggest line, not the only one: storage for model weights, network egress, and the engineering time to keep the server patched, monitored and upgraded. Those don't change with your serving config, so they don't belong in cost per token, but they do belong in the decision.

Your config is part of the price

On the same hardware, serving config moved cost per token from −68.6% to +208% in our runs: batching limits, GPU count and quantization. A self-hosted cost estimate without a measured config is a guess. See vLLM cost per token for the flags that mattered.

Measure before you decide

If you already run a model, measure its cost per token against your own endpoint:

pipx install throttle-pro

throttle check \
  --url http://localhost:8000 \
  --model your-model \
  --gpu-hourly-rate 2.99

It works with any OpenAI-compatible server (vLLM, SGLang, Ollama, LMDeploy), labels the GPU price you typed as an assumption, and measures everything else. If you want this done on your stack with a written report, that's the Cost Audit.

FAQ

Is self-hosting an LLM cheaper than using an API?

It can be when the GPU is busy most of the time. Measure your cost per token under load, divide by your utilization, and compare with the API's per-million price for a comparable model.

How do I calculate the cost of a self-hosted LLM?

Measure output tokens per second at your real concurrency, then: GPU $/hr ÷ (tokens/s × 3600) × 1,000,000 gives $ per million tokens while serving.

What does it cost to run a 72B model?

Measured on one AMD MI300X at $2.99/hr: about $1.67 per million output tokens in BF16 and $1.13 with a pre-quantized FP8 checkpoint, at 32 concurrent requests.

Why is my real self-hosted cost higher than the benchmark?

Idle time. A GPU billed by the hour costs the same when nobody is using it, so low utilization raises the effective cost per token.

Related guides