Spec & Pricing Reference

NVIDIA H200 SXM5 Cloud Pricing & Specs (2026)

The H200 doubles VRAM to 141 GB with 4.8 TB/s HBM3e bandwidth. It runs 70B models on a single GPU and handles 128K+ context windows without the KV-cache memory pressure that constrains the H100.

Target workload: High-concurrency 70B+ LLM inference with extended KV-cache

Memory141GB HBM3e
Bandwidth4.8 TB/s
FP8 TFLOPS1,979
Observed rate$0.00/hr
141 GB VRAM
Manufacturer SpecHIGH
SourceNVIDIA manufacturer specifications
VerifiedSep 30, 2026
Value
141 GB
Methodology

Manufacturer specification for onboard memory.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
4.8 TB/s
Manufacturer SpecHIGH
SourceNVIDIA specification
VerifiedSep 30, 2026
Value
4.8 TB/s GB/s
Methodology

Peak memory bandwidth from the manufacturer specification.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
$0.00/hr on-demand
Observed On-Demand RateHIGH
SourceObserved provider API rate (lowest on-demand row)
VerifiedSep 30, 2026
Value
0 USD/hr
Methodology

Lowest on-demand hourly row for "H200" across tracked providers in data/providers.json; refreshed daily.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
Methodology β†’

NVIDIA H200 SXM5: Key Numbers at a Glance

Rental cost: lowest observed on-demand $0.00/hr, spot rows from $0.00/hr across 4 tracked providers (refreshed daily).

70B fit: Fits Llama 3.3 70B at FP16 (140 GB min VRAM) with KV-cache headroom for 128k context.

Bottleneck: Memory-bandwidth bound across all batch sizes

Observed Pricing for NVIDIA H200 SXM5

Spot and on-demand rows for H200 refreshed from provider APIs. Lowest on-demand: $0.00/hr.

ProviderGPU & VRAMInterconnectSpot Price ($/hr)On-Demand ($/hr)Monthly ($/720h)StatusAction
Lambda Labs
N/A$0/hr
Calculated EstimateMEDIUM
SourceLambda Labs
VerifiedSep 26, 2026
Value
0 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5Γ— spot (conservative; actual range 1.8×–3.2Γ—).

Assumptions & Parameters
  • range: 1.8×–3.2Γ—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$0 / hr$0 / moInstant
Vast.ai
N/A$2.5/hr
Calculated EstimateMEDIUM
SourceVast.ai
VerifiedSep 26, 2026
Value
2.5 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5Γ— spot (conservative; actual range 1.8×–3.2Γ—).

Assumptions & Parameters
  • range: 1.8×–3.2Γ—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$3.4 / hr$0 / moInstant
Spheron
N/A$3.36/hr
Calculated EstimateMEDIUM
SourceSpheron
VerifiedSep 26, 2026
Value
3.36 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5Γ— spot (conservative; actual range 1.8×–3.2Γ—).

Assumptions & Parameters
  • range: 1.8×–3.2Γ—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$4.8 / hr$0 / moInstant
RunPod
N/A$3.89/hr
Calculated EstimateMEDIUM
SourceRunPod
VerifiedSep 26, 2026
Value
3.89 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5Γ— spot (conservative; actual range 1.8×–3.2Γ—).

Assumptions & Parameters
  • range: 1.8×–3.2Γ—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$3.89 / hr$0 / moInstant
Monthly estimates use continuous spot rate Γ— 720h. On-demand calculated at 2.5Γ— spot (range: 1.8×–3.2Γ—). All rates subject to preemption and provider availability.
Data Freshness: Public Cloud APIs & Market Scraping | Refreshed Daily (UTC)Benchmark Baseline: Ubuntu 24.04, CUDA 12.4, vLLM v0.6.x, PagedAttention v2, FlashAttention-3

Compatible Models for NVIDIA H200 SXM5

Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 141 GB. FP16 shows where full precision also fits single-GPU.

Browse all model VRAM pages β†’

Specifications

ArchitectureHopper GH200 β€” 4nm TSMC
Memory141GB HBM3e
Bandwidth4.8 TB/s
InterconnectNVLink 4.0 (900 GB/s)
TDP700W
FP16 TFLOPS989
FP8 TFLOPS1,979
FP4 TFLOPSN/A
Recommended quantizationFP8 / FP4 native
Best cluster topology8-way HGX Baseboard with NVLink 4.0 Mesh

H200 SXM5 141 GB HBM3e KV-Cache Expansion Analysis: The H200's 141 GB HBM3e at 4.8 TB/s enables a critical advantage over the H100: 128k context windows for Llama 70B FP8 on a single GPU. At FP8, the 70B model occupies ~72 GB, leaving 69 GB for KV-cache. At 128k context with batch=1, the KV-cache consumes ~32 GB β€” well within the 69 GB budget. On 2x H100s (TP=2), the same workload requires splitting KV-cache across GPUs, introducing NVLink latency for every attention head. The H200 eliminates this: 1x H200 serving 128k context delivers 15-20% lower time-to-first-token than 2x H100s because there are no cross-GPU KV-cache lookups. For high-concurrency serving (batchβ‰₯32), the 4.8 TB/s bandwidth sustains 2-3x higher throughput before hitting memory stalls. The tradeoff: H200 pricing is typically 15-25% higher than H100, but the single-GPU deployment eliminates tensor parallelism complexity.

Next steps

Compare Alternatives

Head-to-head comparisons against this GPU β€” specs, observed hourly rates, and the workload verdict.

Related GPUs

Frequently Asked Questions

How much does it cost to rent NVIDIA H200 SXM5 per hour?β–Ύ
The lowest observed on-demand rate for NVIDIA H200 SXM5 is $0.00/hr (4 tracked providers); observed spot rows start at $0.00/hr. Rates refresh daily from provider APIs.
Can NVIDIA H200 SXM5 run 70B parameter LLMs?β–Ύ
Fits Llama 3.3 70B at FP16 (140 GB min VRAM) with KV-cache headroom for 128k context.
What is NVIDIA H200 SXM5's memory bandwidth?β–Ύ
NVIDIA H200 SXM5 has 4.8 TB/s of memory bandwidth across 141GB HBM3e. Token generation is memory-bandwidth bound at batch size 1, so decode throughput scales with this figure β€” no tokens/sec value is claimed without a benchmark row.
Rates observed across 4 providers; refreshed daily (UTC).Spec sources: NVIDIA manufacturer specifications.Methodology β†’