⚑Under $0.50/hr🧠VRAM Estimatorβš–Compare GPUs🎁Free LLM APIs🎯Model Index
GPU Price Tracker

NVIDIA H200 SXM5 Cloud Pricing & Rental Rates

The H200 doubles VRAM to 141 GB with 4.8 TB/s HBM3e bandwidth. It runs 70B models on a single GPU and handles 128K+ context windows without the KV-cache memory pressure that constrains the H100.

Target Workload: High-concurrency 70B+ LLM inference with extended KV-cache

Memory141GB HBM3e
Bandwidth4.8 TB/s
FP8 TFLOPS1,979
InterconnectNVLink 4.0 (900 GB/s)

Live Cloud Pricing for NVIDIA H200 SXM5

Spot and reserved rates refreshed from provider APIs. Click a provider to deploy.

ProviderGPU & VRAMInterconnectSpot RateOn-DemandMonthlyStatusAction
Community
NVLink 4.0 (900 GB/s)$2.79 / hr$6.98 / hr$1,707 / moInstant
Bare Metal
NVLink 4.0 (900 GB/s)$3.19 / hr$7.98 / hr$1,952 / moInstant
Dedicated
NVLink 4.0 (900 GB/s)$3.99 / hr$9.98 / hr$2,442 / moInstant
Cloud
NVLink 4.0 (900 GB/s)$4.31 / hr$10.77 / hr$2,638 / moInstant
Data Freshness: Public Cloud APIs & Market Scraping | Refreshed Daily (UTC)Benchmark Baseline: Ubuntu 24.04, CUDA 12.4, vLLM v0.6.x, PagedAttention v2, FlashAttention-3

Models That Fit NVIDIA H200 SXM5 (141GB HBM3e)

Deterministic VRAM calculation (FP8) via entity graph. Only editorial and enriched models shown.

MODELPARAMSCONTEXTFP8 VRAMFIT?CALCULATOR
DeepSeek R1 Distill 70B70B128K93.9 GBβœ“ fitsPre-filled β†’
DeepSeek R1 Distill Qwen 32B32B128K42.9 GBβœ“ fitsPre-filled β†’
Llama 3.3 70B Instruct70.6B128K94.7 GBβœ“ fitsPre-filled β†’
Llama 3.1 8B Instruct8.03B128K10.8 GBβœ“ fitsPre-filled β†’
Llama 3.2 3B Instruct3.2B128K4.3 GBβœ“ fitsPre-filled β†’
Llama 3.2 1B Instruct1B128K1.3 GBβœ“ fitsPre-filled β†’
Qwen 2.5 Coder 32B32.5B128K43.6 GBβœ“ fitsPre-filled β†’
Qwen 2.5 Coder 14B14.7B128K19.7 GBβœ“ fitsPre-filled β†’
Qwen 2.5 Coder 7B7.61B128K10.2 GBβœ“ fitsPre-filled β†’
Qwen 2.5 72B Instruct72.7B128K97.5 GBβœ“ fitsPre-filled β†’
Qwen 2.5 14B Instruct14.7B128K19.7 GBβœ“ fitsPre-filled β†’
Qwen 2.5 7B Instruct7.61B128K10.2 GBβœ“ fitsPre-filled β†’

Inference & Serving Capacity

Practical model feasibility, max batch sizes, and KV-cache retention limits for NVIDIA H200 SXM5.

Llama 3.3 70B

FEASIBLE
Max Batch Size:16-32
KV-Cache:~32 GB at 128k ctx

Single-GPU FP8, full 128k context

DeepSeek 671B

FEASIBLE
Max Batch Size:2-4
KV-Cache:Large KV-cache budget

141 GB fits 16 experts per GPU, 4-GPU cluster

Qwen 2.5 32B

FEASIBLE
Max Batch Size:64-128
KV-Cache:Massive headroom

Single GPU handles full model + large batches

vLLM Throughput (FP8)

~135 tok/s (vLLM, Llama 70B FP8, batch=1)

Estimated tokens/second, single GPU, Llama-class model

Max Context Window (Llama 70B)

128k+ tokens (FP8) β€” KV-cache fits comfortably with batch headroom

Maximum context length before KV-cache eviction

Data Freshness: Verified via Public Cloud APIs & Market Scraping | Refreshed Daily (UTC)Benchmark Testing Baseline: Ubuntu 24.04, CUDA 12.4, vLLM v0.6.x (PagedAttention v2, FlashAttention-3), BF16/FP8 weights, Batch Size = 1 unless specified.Methodology β†’

Hardware Bottleneck Analysis

Whether NVIDIA H200 SXM5 is compute-bound (TFLOPS) or memory-bandwidth bound (GB/s) across workloads.

Bottleneck Classification

Memory-bandwidth bound across all batch sizes

Recommended Quantization

FP8 / FP4 native

Best Cluster Topology

8-way HGX Baseboard with NVLink 4.0 Mesh

Deep Analysis

H200 SXM5 141 GB HBM3e KV-Cache Expansion Analysis: The H200's 141 GB HBM3e at 4.8 TB/s enables a critical advantage over the H100: 128k context windows for Llama 70B FP8 on a single GPU. At FP8, the 70B model occupies ~72 GB, leaving 69 GB for KV-cache. At 128k context with batch=1, the KV-cache consumes ~32 GB β€” well within the 69 GB budget. On 2x H100s (TP=2), the same workload requires splitting KV-cache across GPUs, introducing NVLink latency for every attention head. The H200 eliminates this: 1x H200 serving 128k context delivers 15-20% lower time-to-first-token than 2x H100s because there are no cross-GPU KV-cache lookups. For high-concurrency serving (batchβ‰₯32), the 4.8 TB/s bandwidth sustains 2-3x higher throughput before hitting memory stalls. The tradeoff: H200 pricing is typically 15-25% higher than H100, but the single-GPU deployment eliminates tensor parallelism complexity.

Architecture & Die Breakdown

Architecture

Hopper GH200 β€” 4nm TSMC

TDP

700W

Memory Subsystem

141GB HBM3e at 4.8 TB/s bandwidth. High Bandwidth Memory provides the throughput needed to keep tensor cores fed during large batch inference.

Interconnect

NVLink 4.0 (900 GB/s). Enables multi-GPU tensor parallelism with high-bandwidth, low-latency GPU-to-GPU communication.

Precision Performance

PrecisionTFLOPSUse Case
FP81,979Training & inference with mixed-precision
FP16989Full-precision training, fine-tuning, evaluation

Break-Even ROI Calculator

Monthly hours where reserved pricing beats spot for NVIDIA H200 SXM5. Above the break-even point, reserve commits save money.

Spheron

Spot Rate:$2.79/hr
Reserved Rate:$2.37/hr
Break-Even:340 hrs/mo
Monthly Savings:$301/mo

RunPod

Spot Rate:$6.98/hr
Reserved Rate:$5.93/hr
Break-Even:340 hrs/mo
Monthly Savings:$754/mo

Lambda Labs

Spot Rate:$2.99/hr
Reserved Rate:$2.54/hr
Break-Even:340 hrs/mo
Monthly Savings:$323/mo

Break-even at 340 hours/month: if you run NVIDIA H200 SXM5 more than 340 hours per month, reserved pricing on all three providers saves money. At 720 hours/month (24/7), reserved saves $754+/mo vs on-demand.

Related GPUs

Frequently Asked Questions

How much does it cost to rent an NVIDIA H200 SXM5 per hour?β–Ύ
The lowest spot rate for NVIDIA H200 SXM5 is $2.79/hr. On-demand pricing starts at $6.98/hr depending on the provider and region. Reserved commitments can reduce costs by ~15%.
What is the monthly reserved pricing for NVIDIA H200 SXM5?β–Ύ
Monthly reserved pricing for NVIDIA H200 SXM5 ranges from $4,272/mo to $5,026/mo depending on commitment level and provider.
Can an NVIDIA H200 SXM5 run 70B parameter LLMs?β–Ύ
NVIDIA H200 SXM5 with 141GB HBM3e VRAM can run 70B parameter models with tensor parallelism across 1-2 GPUs. At FP8 precision, the model fits within 141GB HBM3e with room for KV-cache.
Is spot pricing reliable for distributed NVIDIA H200 SXM5 training?β–Ύ
Spot pricing offers 30-60% savings but carries preemption risk. For distributed training, implement SIGTERM handlers with S3 checkpointing every 30 minutes. Vast.ai and RunPod provide 30-second eviction warnings. For critical workloads, reserved pricing eliminates preemption entirely.

Compare alternatives & next steps

Hardware Analysis & Practical Guides

What should I do next?