Tensor parallel size and cost: one GPU vs two for a 72B model
Splitting a model across two GPUs makes it faster. It almost never makes it cheaper. Here is a clean measurement of how much more you pay.
The 1:1
Qwen2.5-72B-Instruct in BF16, vLLM 0.23.1 on ROCm, the same weights and the same workload (32 concurrent requests, 256 max output tokens, cold cache). The only change was --tensor-parallel-size: 1 GPU at $2.99/hr, then 2 GPUs at $5.98/hr. Four checks on one GPU, three on two.
vllm serve Qwen/Qwen2.5-72B-Instruct --tensor-parallel-size 1
vllm serve Qwen/Qwen2.5-72B-Instruct --tensor-parallel-size 2
Results
| 1× MI300X | 2× MI300X | |
|---|---|---|
| Price | $2.99/hr | $5.98/hr |
| Time per batch of 32 requests | 12.7 s | 8.1 s |
| Output tokens per second | ~498 | ~777 |
| $ per million output tokens | $1.67 | $2.14 (+28%) |
| Answers | Same length per request; 5 of 6 spot-check answers identical, the 6th reworded | |
Run-to-run noise was about 1% and the confidence intervals don't overlap, so this is a real difference, in all three checks.
Why two GPUs cost more per token
Two GPUs doubled the hourly price but only raised throughput 1.56×. Tensor parallelism splits every layer across GPUs, and they have to exchange activations at each step, so you never get the full 2×. Cost per token is price divided by throughput, so 2 ÷ 1.56 ≈ 1.28: 28% more per token.
That doesn't make tensor parallel wrong. It buys latency: each batch of 32 requests finished about 36% sooner. If you need that speed for users, it is worth paying for. If you are serving batch jobs or have latency headroom, one GPU per replica (and more replicas for more traffic) is the cheaper layout, as long as the model fits.
Choosing a layout for your traffic
Work it out from three numbers: how many requests you need to serve at peak, the latency you promised, and cost per token for each layout at that load. If one GPU per replica meets your latency target, adding replicas scales throughput at the single-GPU price per token. If it doesn't, compare the tensor parallel premium (28% here) with what the extra speed is worth to you. Either way, measure at your real concurrency: the gap between layouts changes with load, so a number taken at 32 concurrent requests may not hold at 4 or 128.
When the model has to be split
A 72B model in BF16 needs roughly 145 GB for its weights alone (72.7B parameters × 2 bytes), which is why it fits on one 192 GB MI300X but not on one 80 GB GPU. If your model doesn't fit, the choice is between tensor parallel and quantization. Pre-quantized FP8 brought the same 72B to $1.13 per million tokens on a single MI300X.
Measure your own split
Pass the total hourly price of every GPU the server uses:
throttle check --url http://localhost:8000 \
--model Qwen/Qwen2.5-72B-Instruct \
--gpu-hourly-rate 5.98 --label tp2 --config gpus=2
Compare the printed $ per million tokens between the two setups. In Throttle 0.5.1 the verdict line judges throughput at the baseline's hourly rate, so for a GPU-count change read the dollar figures; the next release counts the extra GPUs as a real cost in the verdict too.
FAQ
Does tensor parallel size 2 reduce cost per token?
Usually not. In a 1:1 on AMD MI300X, Qwen2.5-72B cost $1.67 per million output tokens on one GPU and $2.14 on two: 1.56× faster, 28% more per token.
When should I use tensor parallelism?
When the model doesn't fit on one GPU, or when you need lower latency per request and will pay for it.
How do I compare cost across different GPU counts?
Use the total hourly price of all GPUs in the server and compare dollars per million tokens at the same workload and concurrency.
Does tensor parallel change the model's answers?
It shouldn't change them meaningfully. In our run, output length per request was the same and 5 of 6 spot-check answers were identical.
Related guides
vLLM cost per token: measure it, then change one flag at a time
Measure vLLM cost per million tokens on your own GPU, then test max-num-seqs, tensor parallel size and FP8 one at a time. Real results from AMD MI300X and NVIDIA A100 runs.
CalculatorLLM cost per million tokens: the formula, a calculator, and real numbers
How to calculate LLM inference cost per million tokens from your GPU price and throughput, with a calculator, a worked example from a real AMD MI300X run, and the pitfalls that make the number wrong.