NVIDIA H200 SXM5 Cloud Pricing & Rental Rates
The H200 doubles VRAM to 141 GB with 4.8 TB/s HBM3e bandwidth. It runs 70B models on a single GPU and handles 128K+ context windows without the KV-cache memory pressure that constrains the H100.
Target Workload: High-concurrency 70B+ LLM inference with extended KV-cache
Live Cloud Pricing for NVIDIA H200 SXM5
Spot and reserved rates refreshed from provider APIs. Click a provider to deploy.
| Provider | GPU & VRAM | Interconnect | Spot Rate | On-Demand | Monthly | Status | Action | |
|---|---|---|---|---|---|---|---|---|
Community | NVLink 4.0 (900 GB/s) | $2.79 / hr | $6.98 / hr | $1,707 / mo | Instant | |||
Bare Metal | NVLink 4.0 (900 GB/s) | $3.19 / hr | $7.98 / hr | $1,952 / mo | Instant | |||
Dedicated | NVLink 4.0 (900 GB/s) | $3.99 / hr | $9.98 / hr | $2,442 / mo | Instant | |||
Cloud | NVLink 4.0 (900 GB/s) | $4.31 / hr | $10.77 / hr | $2,638 / mo | Instant |
Models That Fit NVIDIA H200 SXM5 (141GB HBM3e)
Deterministic VRAM calculation (FP8) via entity graph. Only editorial and enriched models shown.
| MODEL | PARAMS | CONTEXT | FP8 VRAM | FIT? | CALCULATOR |
|---|---|---|---|---|---|
| DeepSeek R1 Distill 70B | 70B | 128K | 93.9 GB | β fits | Pre-filled β |
| DeepSeek R1 Distill Qwen 32B | 32B | 128K | 42.9 GB | β fits | Pre-filled β |
| Llama 3.3 70B Instruct | 70.6B | 128K | 94.7 GB | β fits | Pre-filled β |
| Llama 3.1 8B Instruct | 8.03B | 128K | 10.8 GB | β fits | Pre-filled β |
| Llama 3.2 3B Instruct | 3.2B | 128K | 4.3 GB | β fits | Pre-filled β |
| Llama 3.2 1B Instruct | 1B | 128K | 1.3 GB | β fits | Pre-filled β |
| Qwen 2.5 Coder 32B | 32.5B | 128K | 43.6 GB | β fits | Pre-filled β |
| Qwen 2.5 Coder 14B | 14.7B | 128K | 19.7 GB | β fits | Pre-filled β |
| Qwen 2.5 Coder 7B | 7.61B | 128K | 10.2 GB | β fits | Pre-filled β |
| Qwen 2.5 72B Instruct | 72.7B | 128K | 97.5 GB | β fits | Pre-filled β |
| Qwen 2.5 14B Instruct | 14.7B | 128K | 19.7 GB | β fits | Pre-filled β |
| Qwen 2.5 7B Instruct | 7.61B | 128K | 10.2 GB | β fits | Pre-filled β |
Inference & Serving Capacity
Practical model feasibility, max batch sizes, and KV-cache retention limits for NVIDIA H200 SXM5.
Llama 3.3 70B
FEASIBLESingle-GPU FP8, full 128k context
DeepSeek 671B
FEASIBLE141 GB fits 16 experts per GPU, 4-GPU cluster
Qwen 2.5 32B
FEASIBLESingle GPU handles full model + large batches
vLLM Throughput (FP8)
~135 tok/s (vLLM, Llama 70B FP8, batch=1)
Estimated tokens/second, single GPU, Llama-class model
Max Context Window (Llama 70B)
128k+ tokens (FP8) β KV-cache fits comfortably with batch headroom
Maximum context length before KV-cache eviction
Hardware Bottleneck Analysis
Whether NVIDIA H200 SXM5 is compute-bound (TFLOPS) or memory-bandwidth bound (GB/s) across workloads.
Bottleneck Classification
Memory-bandwidth bound across all batch sizes
Recommended Quantization
FP8 / FP4 native
Best Cluster Topology
8-way HGX Baseboard with NVLink 4.0 Mesh
Deep Analysis
H200 SXM5 141 GB HBM3e KV-Cache Expansion Analysis: The H200's 141 GB HBM3e at 4.8 TB/s enables a critical advantage over the H100: 128k context windows for Llama 70B FP8 on a single GPU. At FP8, the 70B model occupies ~72 GB, leaving 69 GB for KV-cache. At 128k context with batch=1, the KV-cache consumes ~32 GB β well within the 69 GB budget. On 2x H100s (TP=2), the same workload requires splitting KV-cache across GPUs, introducing NVLink latency for every attention head. The H200 eliminates this: 1x H200 serving 128k context delivers 15-20% lower time-to-first-token than 2x H100s because there are no cross-GPU KV-cache lookups. For high-concurrency serving (batchβ₯32), the 4.8 TB/s bandwidth sustains 2-3x higher throughput before hitting memory stalls. The tradeoff: H200 pricing is typically 15-25% higher than H100, but the single-GPU deployment eliminates tensor parallelism complexity.
Architecture & Die Breakdown
Architecture
Hopper GH200 β 4nm TSMC
TDP
700W
Memory Subsystem
141GB HBM3e at 4.8 TB/s bandwidth. High Bandwidth Memory provides the throughput needed to keep tensor cores fed during large batch inference.
Interconnect
NVLink 4.0 (900 GB/s). Enables multi-GPU tensor parallelism with high-bandwidth, low-latency GPU-to-GPU communication.
Precision Performance
| Precision | TFLOPS | Use Case |
|---|---|---|
| FP8 | 1,979 | Training & inference with mixed-precision |
| FP16 | 989 | Full-precision training, fine-tuning, evaluation |
Break-Even ROI Calculator
Monthly hours where reserved pricing beats spot for NVIDIA H200 SXM5. Above the break-even point, reserve commits save money.
Spheron
RunPod
Lambda Labs
Break-even at 340 hours/month: if you run NVIDIA H200 SXM5 more than 340 hours per month, reserved pricing on all three providers saves money. At 720 hours/month (24/7), reserved saves $754+/mo vs on-demand.
Related GPUs
Frequently Asked Questions
How much does it cost to rent an NVIDIA H200 SXM5 per hour?βΎ
What is the monthly reserved pricing for NVIDIA H200 SXM5?βΎ
Can an NVIDIA H200 SXM5 run 70B parameter LLMs?βΎ
Is spot pricing reliable for distributed NVIDIA H200 SXM5 training?βΎ
Compare alternatives & next steps