NVIDIA H200 SXM5 Cloud Pricing & Specs (2026)
The H200 doubles VRAM to 141 GB with 4.8 TB/s HBM3e bandwidth. It runs 70B models on a single GPU and handles 128K+ context windows without the KV-cache memory pressure that constrains the H100.
Target workload: High-concurrency 70B+ LLM inference with extended KV-cache
141 GB VRAM
Manufacturer specification for onboard memory.
4.8 TB/s
Peak memory bandwidth from the manufacturer specification.
$0.00/hr on-demand
Lowest on-demand hourly row for "H200" across tracked providers in data/providers.json; refreshed daily.
NVIDIA H200 SXM5: Key Numbers at a Glance
Rental cost: lowest observed on-demand $0.00/hr, spot rows from $0.00/hr across 4 tracked providers (refreshed daily).
70B fit: Fits Llama 3.3 70B at FP16 (140 GB min VRAM) with KV-cache headroom for 128k context.
Bottleneck: Memory-bandwidth bound across all batch sizes
Observed Pricing for NVIDIA H200 SXM5
Spot and on-demand rows for H200 refreshed from provider APIs. Lowest on-demand: $0.00/hr.
Featured GPU Pricing Pages
80GB HBM3 β’ 3.35 TB/s β’ NVLink 4.0
40GB HBM2 β’ 1.55 TB/s β’ NVLink 3.0
Verified spot rates & reserved pricing
24GB GDDR6 β’ 0.62 TB/s β’ PCIe 4.0
24GB GDDR6X β’ 0.94 TB/s β’ PCIe 4.0
16GB GDDR6 β’ 0.32 TB/s β’ PCIe 3.0
Compatible Models for NVIDIA H200 SXM5
Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 141 GB. FP16 shows where full precision also fits single-GPU.
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
Specifications
H200 SXM5 141 GB HBM3e KV-Cache Expansion Analysis: The H200's 141 GB HBM3e at 4.8 TB/s enables a critical advantage over the H100: 128k context windows for Llama 70B FP8 on a single GPU. At FP8, the 70B model occupies ~72 GB, leaving 69 GB for KV-cache. At 128k context with batch=1, the KV-cache consumes ~32 GB β well within the 69 GB budget. On 2x H100s (TP=2), the same workload requires splitting KV-cache across GPUs, introducing NVLink latency for every attention head. The H200 eliminates this: 1x H200 serving 128k context delivers 15-20% lower time-to-first-token than 2x H100s because there are no cross-GPU KV-cache lookups. For high-concurrency serving (batchβ₯32), the 4.8 TB/s bandwidth sustains 2-3x higher throughput before hitting memory stalls. The tradeoff: H200 pricing is typically 15-25% higher than H100, but the single-GPU deployment eliminates tensor parallelism complexity.
Next steps
Compare Alternatives
Head-to-head comparisons against this GPU β specs, observed hourly rates, and the workload verdict.
The H200 eliminates 2-GPU tensor sharding for Llama 70B at 128k context. The H100 remains the lower-cost option for sub-32k workloads at $1.89/hr vs $2.50/hr.
H100 SXM5 is the cost-efficient workhorse: identical 1,979 FP8 TFLOPS and NVLink 4.0 at $1.89/hr observed β $0.90/hr less than the H200. The H200's 141 GB (+76% VRAM, +43% bandwidth) is the unlock for long-context work: 70B at FP8 with full 128k KV-cache fits one GPU, where the H100's 80 GB forces eviction or shorter context. Choose H100 for cost-efficient standard serving; choose H200 when context length or single-GPU simplicity outruns.