What LLM inference really costs
How to calculate and measure cost per million tokens on your own GPUs, and what real config changes did to it. Every number comes from a measured run.
LLM cost per million tokens: the formula, a calculator, and real numbers
How to calculate LLM inference cost per million tokens from your GPU price and throughput, with a calculator, a worked example from a real AMD MI300X run, and the pitfalls that make the number wrong.
GuidevLLM cost per token: measure it, then change one flag at a time
Measure vLLM cost per million tokens on your own GPU, then test max-num-seqs, tensor parallel size and FP8 one at a time. Real results from AMD MI300X and NVIDIA A100 runs.
Case studyFP8 quantization in vLLM: 47% cheaper tokens, until you read the answers
Real AMD MI300X results for FP8 quantization in vLLM: pre-quantized FP8 cut 72B cost per token by 32% with correct answers, while on-the-fly FP8 looked 41–47% cheaper but produced broken output.
Case studyTensor parallel size and cost: one GPU vs two for a 72B model
Does tensor_parallel_size 2 make tokens cheaper? A measured 1:1 on AMD MI300X: Qwen2.5-72B cost $1.67 per million tokens on one GPU and $2.14 on two, 1.56× faster but 28% more per token.
GuideSelf-hosted LLM cost: how to compare it with an API, honestly
What does a self-hosted LLM really cost? How to turn GPU rental prices into dollars per million tokens, why utilization decides the answer, and real measurements on AMD MI300X.