Notes from real runs.
What self-hosted LLM inference actually costs, why the number moves when nothing changed, and how to tell a real saving from noise. Every figure comes from a recorded run.
I changed nothing and my LLM server got 27% more expensive
Four cost checks of the same untouched Ollama server came back at $8.77, $8.86, $11.26 and $10.02 per million output tokens. Why same-server runs swing, and the noise bound I now use before believing any before/after number.
GuideHow to measure cost per million tokens on a self-hosted LLM (vLLM, SGLang, Ollama)
A practical guide to turning GPU dollars per hour into dollars per million tokens for vLLM, SGLang and Ollama, the four pitfalls that make the number wrong, and how to re-check it after every config change and in CI.
Case studyOne vLLM flag, −68.6% cost per token: what a counterbalanced benchmark looks like
A real A100 run where changing vLLM's max_num_seqs from 1 to 8 took cost from $0.746 to $0.234 per million output tokens, measured with a six-position B‑C‑B / C‑B‑C protocol. How the protocol works, the raw numbers, and what the result does not mean.