These are modeled estimates from vendor-published specs (TDP, HBM capacity, list-price ballparks), not measurements from a testbed we operate. Every assumption is stated below so you can re-run the arithmetic — plug your own quotes into the cluster modeler and export the CSV before any procurement conversation.
Headline numbers (1,024 GPUs · 128k context · PUE 1.4)
| Config | Est. capex | Power draw | Power $/mo | Step latency | Energy / 1k tok |
|---|---|---|---|---|---|
| H100 · IB NDR | $31.7M | 1,362 kW | $119k | 64.8ms | 24.5 Wh |
| H100 · RoCEv2 | $29.9M | 1,362 kW | $119k | 100.8ms | 38.1 Wh |
| H200 · IB NDR | $35.8M | 1,362 kW | $119k | 64.8ms | 24.5 Wh |
| B200 · IB NDR | $44.0M | 1,792 kW | $157k | 64.8ms | 32.3 Wh |
| B200 · RoCEv2 | $42.2M | 1,792 kW | $157k | 100.8ms | 50.2 Wh |
What the model actually says
The fabric decision dominates latency. At 1,024 GPUs, RoCEv2 adds 56% step latency versus InfiniBand NDR (64.8ms → 100.8ms) under identical GPUs and context. The networking line item ($3.0k/GPU for IB vs ~$1.2k/GPU for RoCE) is ~4% of capex and buys back more than half the latency penalty. Price the fabric before haggling over GPU SKUs.
Context length dominates everything else. Same B200 / IB cluster: 8k context → 24.3ms steps; 128k → 64.8ms; 1M → 367.2ms and 182.8 Wh per 1k tokens. If your workload is long-context, KV-cache capacity (HBM per GPU: 80GB H100, 141GB H200, 192GB B200 — all vendor-published) is the binding constraint, not FLOPS.
Scale has a floor. A 64-GPU B200 pod models at $2.75M capex and 53.3ms steps — the logarithmic collectives penalty means the first doubling of fabric hurts most. Pilot at 64, measure, then scale; do not extrapolate linearly from 8-GPU nodes.
Assumptions (challenge all of them)
- GPU TDP: 700W (H100/H200 SXM), 1,000W (B200) — vendor-published thermal design power, not measured draw.
- Host overhead: 2 kW per 8-GPU node. PUE 1.4. Electricity $0.12/kWh, 730 h/month.
- List-price ballparks: $28k H100, $32k H200, $40k B200, plus fabric premium per GPU.
- Step latency: 12ms base scaled by KV-cache pressure (context/64k) and a logarithmic fabric penalty. Per-GPU throughput differences between generations are not modeled — treat latency rows as fabric/context comparisons, not GPU verdicts.
Decision impact
Get IB quotes alongside GPU quotes; size HBM for your p99 context, not your median; re-run with your power tariff before signing. Raw per-config numbers download from the modeler above as CSV.