All guides Case study

Tensor parallel size and cost: one GPU vs two for a 72B model

Splitting a model across two GPUs makes it faster. It almost never makes it cheaper. Here is a clean measurement of how much more you pay.

The 1:1

Qwen2.5-72B-Instruct in BF16, vLLM 0.23.1 on ROCm, the same weights and the same workload (32 concurrent requests, 256 max output tokens, cold cache). The only change was --tensor-parallel-size: 1 GPU at $2.99/hr, then 2 GPUs at $5.98/hr. Four checks on one GPU, three on two.

vllm serve Qwen/Qwen2.5-72B-Instruct --tensor-parallel-size 1
vllm serve Qwen/Qwen2.5-72B-Instruct --tensor-parallel-size 2

Results

1× MI300X2× MI300X
Price$2.99/hr$5.98/hr
Time per batch of 32 requests12.7 s8.1 s
Output tokens per second~498~777
$ per million output tokens$1.67$2.14 (+28%)
AnswersSame length per request; 5 of 6 spot-check answers identical, the 6th reworded

Run-to-run noise was about 1% and the confidence intervals don't overlap, so this is a real difference, in all three checks.

Why two GPUs cost more per token

Two GPUs doubled the hourly price but only raised throughput 1.56×. Tensor parallelism splits every layer across GPUs, and they have to exchange activations at each step, so you never get the full 2×. Cost per token is price divided by throughput, so 2 ÷ 1.56 ≈ 1.28: 28% more per token.

That doesn't make tensor parallel wrong. It buys latency: each batch of 32 requests finished about 36% sooner. If you need that speed for users, it is worth paying for. If you are serving batch jobs or have latency headroom, one GPU per replica (and more replicas for more traffic) is the cheaper layout, as long as the model fits.

Choosing a layout for your traffic

Work it out from three numbers: how many requests you need to serve at peak, the latency you promised, and cost per token for each layout at that load. If one GPU per replica meets your latency target, adding replicas scales throughput at the single-GPU price per token. If it doesn't, compare the tensor parallel premium (28% here) with what the extra speed is worth to you. Either way, measure at your real concurrency: the gap between layouts changes with load, so a number taken at 32 concurrent requests may not hold at 4 or 128.

When the model has to be split

A 72B model in BF16 needs roughly 145 GB for its weights alone (72.7B parameters × 2 bytes), which is why it fits on one 192 GB MI300X but not on one 80 GB GPU. If your model doesn't fit, the choice is between tensor parallel and quantization. Pre-quantized FP8 brought the same 72B to $1.13 per million tokens on a single MI300X.

Measure your own split

Pass the total hourly price of every GPU the server uses:

throttle check --url http://localhost:8000 \
  --model Qwen/Qwen2.5-72B-Instruct \
  --gpu-hourly-rate 5.98 --label tp2 --config gpus=2

Compare the printed $ per million tokens between the two setups. In Throttle 0.5.1 the verdict line judges throughput at the baseline's hourly rate, so for a GPU-count change read the dollar figures; the next release counts the extra GPUs as a real cost in the verdict too.

FAQ

Does tensor parallel size 2 reduce cost per token?

Usually not. In a 1:1 on AMD MI300X, Qwen2.5-72B cost $1.67 per million output tokens on one GPU and $2.14 on two: 1.56× faster, 28% more per token.

When should I use tensor parallelism?

When the model doesn't fit on one GPU, or when you need lower latency per request and will pay for it.

How do I compare cost across different GPU counts?

Use the total hourly price of all GPUs in the server and compare dollars per million tokens at the same workload and concurrency.

Does tensor parallel change the model's answers?

It shouldn't change them meaningfully. In our run, output length per request was the same and 5 of 6 spot-check answers were identical.

Related guides