Skip to main content
The Raw Logs
Benchmark·All benchmarks·2026-09-11

B200 vs H100 cluster economics at 1,024 GPUs: modeled capex, power, and step latency

Same fabric, same context: B200 costs ~39% more capex than H100 for ~32% more power draw — the fabric choice moves latency more than the GPU choice.
HardwareModeled 1,024-GPU cluster · H100 / H200 / B200 SXM · InfiniBand NDR vs RoCEv2
MetricsCapex $29.9M–$44.0M · 1,362–1,792 kW · step latency 24–367ms across contexts
Published2026-09-11
Raw JSON →

These are modeled estimates from vendor-published specs (TDP, HBM capacity, list-price ballparks), not measurements from a testbed we operate. Every assumption is stated below so you can re-run the arithmetic — plug your own quotes into the cluster modeler and export the CSV before any procurement conversation.

Headline numbers (1,024 GPUs · 128k context · PUE 1.4)

ConfigEst. capexPower drawPower $/moStep latencyEnergy / 1k tok
H100 · IB NDR$31.7M1,362 kW$119k64.8ms24.5 Wh
H100 · RoCEv2$29.9M1,362 kW$119k100.8ms38.1 Wh
H200 · IB NDR$35.8M1,362 kW$119k64.8ms24.5 Wh
B200 · IB NDR$44.0M1,792 kW$157k64.8ms32.3 Wh
B200 · RoCEv2$42.2M1,792 kW$157k100.8ms50.2 Wh

What the model actually says

The fabric decision dominates latency. At 1,024 GPUs, RoCEv2 adds 56% step latency versus InfiniBand NDR (64.8ms → 100.8ms) under identical GPUs and context. The networking line item ($3.0k/GPU for IB vs ~$1.2k/GPU for RoCE) is ~4% of capex and buys back more than half the latency penalty. Price the fabric before haggling over GPU SKUs.

Context length dominates everything else. Same B200 / IB cluster: 8k context → 24.3ms steps; 128k → 64.8ms; 1M → 367.2ms and 182.8 Wh per 1k tokens. If your workload is long-context, KV-cache capacity (HBM per GPU: 80GB H100, 141GB H200, 192GB B200 — all vendor-published) is the binding constraint, not FLOPS.

Scale has a floor. A 64-GPU B200 pod models at $2.75M capex and 53.3ms steps — the logarithmic collectives penalty means the first doubling of fabric hurts most. Pilot at 64, measure, then scale; do not extrapolate linearly from 8-GPU nodes.

Assumptions (challenge all of them)

  • GPU TDP: 700W (H100/H200 SXM), 1,000W (B200) — vendor-published thermal design power, not measured draw.
  • Host overhead: 2 kW per 8-GPU node. PUE 1.4. Electricity $0.12/kWh, 730 h/month.
  • List-price ballparks: $28k H100, $32k H200, $40k B200, plus fabric premium per GPU.
  • Step latency: 12ms base scaled by KV-cache pressure (context/64k) and a logarithmic fabric penalty. Per-GPU throughput differences between generations are not modeled — treat latency rows as fabric/context comparisons, not GPU verdicts.

Decision impact

Get IB quotes alongside GPU quotes; size HBM for your p99 context, not your median; re-run with your power tariff before signing. Raw per-config numbers download from the modeler above as CSV.