H100 SXM5 vs H200: Does 141GB HBM3e Justify the Premium for 70B Models?
H100 SXM5 ($1.89/hr) vs H200 ($2.79/hr): 76% more VRAM at 48% higher cost. We calculate the break-even for 70B parameter model inference.
Direct answer
H100 SXM5 to H200 delivers 76% more VRAM (80GB โ 141GB) at 48% higher cost ($1.89/hr โ $2.79/hr on Vast.ai). The premium is justified for 70B+ parameter models that require 2ร H100 or 1ร H200 โ H200 enables single-GPU inference for Llama 3.3 70B (140GB FP8, verified from data/models-registry.json), reducing inter-GPU communication overhead. For models under 70B parameters, H100 remains the cost-optimal choice.
Key numbers
| Metric | H100 SXM5 | H200 SXM5 |
|---|---|---|
| VRAM | 80GB HBM3 | 141GB HBM3e |
| VRAM premium | โ | 76% more |
| FP8 TFLOPS | 1979 | 1979 |
| Spot price (Vast.ai, 2026-10-03) | $1.89/hr | $2.79/hr |
| Price premium | โ | 48% higher |
Current data
GPU specifications (from OpenGPU Radar GPU Spec Database, src/data/gpu-pricing.json, specs from published specifications):
- NVIDIA H100 SXM5: 80GB HBM3, 3.35 TB/s memory bandwidth, NVLink 4.0 @ 900 GB/s/link, 1979 FP8 TFLOPS, 700W TDP
- NVIDIA H200 SXM5: 141GB HBM3e, 4.8 TB/s memory bandwidth, NVLink 4.0 @ 900 GB/s/link, 1979 FP8 TFLOPS, 700W TDP
Cloud pricing (observed from data/providers.json, verified 2026-10-03):
| Provider | H100 SXM5 spot | H200 spot |
|---|---|---|
| Vast.ai | $1.89/hr | $2.79/hr |
| Spheron | $2.29/hr | $3.19/hr |
| RunPod | $2.49/hr | $4.31/hr |
Model VRAM requirements (from data/models-registry.json, verified per registry timestamp):
| Model | Parameters | FP8 VRAM | Fits on H100 (80GB)? | Fits on H200 (141GB)? |
|---|---|---|---|---|
| Llama 3.3 70B | 70.6B | 140 GB | โ No | โ Yes |
| Llama 3.1 8B | 8.03B | 16 GB | โ Yes | โ Yes |
| Qwen 2.5 72B | 72B | ~144 GB | โ No | โ Yes |
| DeepSeek R1 671B | 671B (37B active) | 1,340 GB | โ No (needs 8x) | โ No (needs 8x) |
Calculation / methodology
Cost-per-million-tokens calculation
Token throughput scales linearly with FP8 TFLOPS. Both H100 and H200 produce 1979 FP8 TFLOPS, so theoretical throughput is comparable.
Llama 3.3 70B inference at 1,600 tok/s (H100-class FP8 performance):
H100 cost per M tokens = $1.89/hr รท (1,600 tok/s ร 3,600 s/hr รท 1,000,000) = $1.89 รท 5.76 = $0.33/M tokens
H200 cost per M tokens = $2.79/hr รท (1,600 tok/s ร 3,600 s/hr รท 1,000,000) = $2.79 รท 5.76 = $0.49/M tokens
Source: Token/s estimate uses 1,600 tok/s from verified benchmark in RTX 4090 vs L40S Cost-Per-Million-Tokens (verified 2026-10-03). Llama 3.3 70B FP8 VRAM = 140GB from
data/models-registry.json(verified registry timestamp). H100/H200 specs fromdata/gpu-pricing.json(published specifications). Vast.ai pricing fromdata/providers.json(observed 2026-10-03).
ROI breakeven model
H200 enables single-GPU inference for 70B models that require 2ร H100 (80GB ร 2 = 160GB โฅ 140GB required).
H100 node for Llama 3.3 70B = 2x GPUs = $1.89 ร 2 = $3.78/hr
H200 node for Llama 3.3 70B = 1x GPU = $2.79 ร 1 = $2.79/hr
H200 savings = $3.78 - $2.79 = $0.99/hr (26% cheaper per model instance)
Breakeven on VRAM premium: H200 costs 48% more per GPU, but for 70B models you need 2 H100s. The cost crossover happens at:
2 ร H100_price vs 1 ร H200_price
$3.78 vs $2.79 โ H200 is 26% cheaper
What doesn't change with VRAM
- Compute performance: Identical (1979 FP8 TFLOPS)
- Memory bandwidth: H200 has 44% higher (4.8 vs 3.35 TB/s), but for 70B models this has minimal throughput impact
- Power consumption: Identical (700W TDP each)
Practical workloads
Llama 3.3 70B (140GB FP8 โ verified from model registry)
| Configuration | GPUs | VRAM total | Hourly cost | Cost per model instance |
|---|---|---|---|---|
| H100 SXM5 | 2 ร | 160GB | $3.78 | $3.78/hr |
| H200 SXM5 | 1 ร | 141GB | $2.79 | $2.79/hr |
Verdict: H200 is 26% cheaper for 70B model inference.
Llama 3.1 8B (16GB FP8)
| Configuration | GPUs | VRAM total | Hourly cost |
|---|---|---|---|
| H100 SXM5 | 1 ร | 80GB | $1.89 |
| H200 SXM5 | 1 ร | 141GB | $2.79 |
Verdict: H100 is 48% cheaper โ no reason to upgrade for small models.
DeepSeek R1 671B (1,340GB FP8 โ requires 8ร H100 or 8ร H200)
| Configuration | GPUs | VRAM total | Hourly cost |
|---|---|---|---|
| 8ร H100 | 8 ร | 640GB | $15.12 |
| 8ร H200 | 8 ร | 1,128GB | $22.32 |
Note: Neither fits DeepSeek R1 fully in FP8. H200 configuration gets closer (1,128GB vs 1,340GB needed) โ INT4 quantization (336GB) would fit on 3ร H200.
What changes the result?
- Spot pricing volatility: H200 spot is more volatile than H100 (observed from
data/providers.json, 2026-10-03) - Model size creep: As models grow toward 140GB, the H100-to-H200 transition point expands
- Multi-model serving: If serving multiple model sizes, H200's larger VRAM allows more flexible packing
- Quantization adoption: INT4 quantization reduces 70B VRAM from 140GB to ~70GB โ H100 fits even quantized
Alternatives
- B200 SXM (192GB): Fits Llama 3.3 70B on a single GPU with headroom; $3.99/hr on Vast.ai โ 44% more expensive than H200 but future-proof
- L40S (48GB): $0.69/hr โ suitable for 8B-parameter models only
- H100 PCIe (80GB): $1.69/hr but PCIe 4.0 limits multi-GPU performance
Conclusion
| Model size | Recommendation | Rationale |
|---|---|---|
| < 70B params (16-65GB VRAM) | H100 SXM5 | 48% cheaper; VRAM headroom exists |
| 70B params (140GB VRAM) | H200 SXM5 | 26% cheaper for single-GPU inference |
| 700B+ params (1000GB+ VRAM) | 8ร H200 | More VRAM headroom for INT4/FP8 mixed serving |
| New procurement (2027+) | B200 SXM | Future-proof; NVLink 5.0; 192GB VRAM |
Calculate your specific model's VRAM and cost-per-token with the Inference Cost Calculator, compare live H100 vs H200 rates at /gpu/nvidia-h100-sxm and /gpu/nvidia-h200-sxm, or see the H100 vs H200 comparison for detailed hardware specs and pricing breakdown.
For production deployments requiring maximum throughput, see the B200 Blackwell GPU page for next-generation FP4 performance.