⚑Under $0.50/hr🧠VRAM Estimatorβš–Compare GPUs🎁Free LLM APIs🎯Model Index
GPU Price Tracker

NVIDIA H100 SXM5 Cloud Pricing & Rental Rates

The H100 SXM5 is the workhorse of AI training clusters. Its 80 GB HBM3 memory handles 70B parameter models with tensor parallelism across 2-4 GPUs, while 3.35 TB/s bandwidth keeps the compute units fed during all-reduce operations.

Target Workload: Multi-node 70B+ LLM pre-training & fine-tuning

Memory80GB HBM3
Bandwidth3.35 TB/s
FP8 TFLOPS1,979
InterconnectNVLink 4.0 (900 GB/s)

Live Cloud Pricing for NVIDIA H100 SXM5

Spot and reserved rates refreshed from provider APIs. Click a provider to deploy.

ProviderGPU & VRAMInterconnectSpot RateOn-DemandMonthlyStatusAction
Community
NVLink 4.0 (900 GB/s)$1.89 / hr$4.72 / hr$1,157 / moInstant
Bare Metal
NVLink 4.0 (900 GB/s)$2.29 / hr$5.73 / hr$1,401 / moInstant
Dedicated
NVLink 4.0 (900 GB/s)$2.99 / hr$7.48 / hr$1,830 / moInstant
Cloud
NVLink 4.0 (900 GB/s)$3.49 / hr$8.73 / hr$2,136 / moInstant
Data Freshness: Public Cloud APIs & Market Scraping | Refreshed Daily (UTC)Benchmark Baseline: Ubuntu 24.04, CUDA 12.4, vLLM v0.6.x, PagedAttention v2, FlashAttention-3

Models That Fit NVIDIA H100 SXM5 (80GB HBM3)

Deterministic VRAM calculation (FP8) via entity graph. Only editorial and enriched models shown.

MODELPARAMSCONTEXTFP8 VRAMFIT?CALCULATOR
DeepSeek R1 Distill Qwen 32B32B128K42.9 GBβœ“ fitsPre-filled β†’
Llama 3.1 8B Instruct8.03B128K10.8 GBβœ“ fitsPre-filled β†’
Llama 3.2 3B Instruct3.2B128K4.3 GBβœ“ fitsPre-filled β†’
Llama 3.2 1B Instruct1B128K1.3 GBβœ“ fitsPre-filled β†’
Qwen 2.5 Coder 32B32.5B128K43.6 GBβœ“ fitsPre-filled β†’
Qwen 2.5 Coder 14B14.7B128K19.7 GBβœ“ fitsPre-filled β†’
Qwen 2.5 Coder 7B7.61B128K10.2 GBβœ“ fitsPre-filled β†’
Qwen 2.5 14B Instruct14.7B128K19.7 GBβœ“ fitsPre-filled β†’
Qwen 2.5 7B Instruct7.61B128K10.2 GBβœ“ fitsPre-filled β†’
Mistral NeMo 12B12B128K16.1 GBβœ“ fitsPre-filled β†’
Gemma 2 27B27B8K35.7 GBβœ“ fitsPre-filled β†’
Gemma 2 9B9.24B8K12.2 GBβœ“ fitsPre-filled β†’

Inference & Serving Capacity

Practical model feasibility, max batch sizes, and KV-cache retention limits for NVIDIA H100 SXM5.

Llama 3.3 70B

FEASIBLE
Max Batch Size:4-8
KV-Cache:~16 GB at 32k ctx

FP8 native, 2-GPU tensor parallelism recommended

DeepSeek 671B

FEASIBLE
Max Batch Size:1-2
KV-Cache:Expert parallelism required

8-way NVLink mesh, 80 GB per GPU fits 8 experts

Qwen 2.5 32B

FEASIBLE
Max Batch Size:32-64
KV-Cache:Plenty of headroom

FP8 on single GPU, high batch throughput

vLLM Throughput (FP8)

~120 tok/s (vLLM, Llama 70B FP8, batch=1)

Estimated tokens/second, single GPU, Llama-class model

Max Context Window (Llama 70B)

32k tokens (FP8) β€” KV-cache consumes ~16 GB at batch=1

Maximum context length before KV-cache eviction

Data Freshness: Verified via Public Cloud APIs & Market Scraping | Refreshed Daily (UTC)Benchmark Testing Baseline: Ubuntu 24.04, CUDA 12.4, vLLM v0.6.x (PagedAttention v2, FlashAttention-3), BF16/FP8 weights, Batch Size = 1 unless specified.Methodology β†’

Hardware Bottleneck Analysis

Whether NVIDIA H100 SXM5 is compute-bound (TFLOPS) or memory-bandwidth bound (GB/s) across workloads.

Bottleneck Classification

Compute-bound at small batch; memory-bandwidth bound at batchβ‰₯64

Recommended Quantization

FP8 / FP4 native

Best Cluster Topology

8-way HGX Baseboard with NVLink 4.0 Mesh

Deep Analysis

H100 SXM5 8-Way HGX AllReduce Analysis: In an 8-way HGX baseboard, each GPU connects via NVLink 4.0 at 900 GB/s bidirectional bandwidth. For Llama 70B FP8 tensor parallelism, TP=2 splits the 72 GB model across 2 GPUs (36 GB each) with 16 GB headroom for KV-cache at 32k context. TP=4 splits across 4 GPUs (18 GB each) but introduces 3 all-reduce sync points per forward pass β€” the 900 GB/s NVLink mesh absorbs this with <5% communication overhead. At TP=8, the model shards to 9 GB per GPU, but NCCL all-reduce latency increases to ~8% due to the ring-all-reduce topology across 8 nodes. The optimal config is TP=2 for inference (lowest latency) and TP=4 for training (better compute utilization). NVLink saturation occurs at batchβ‰₯128 when the 3.35 TB/s HBM3 bandwidth cannot sustain the 1,979 FP8 TFLOPS β€” this is the fundamental ceiling for single-node throughput.

Architecture & Die Breakdown

Architecture

Hopper GH100 β€” 4nm TSMC

TDP

700W

Memory Subsystem

80GB HBM3 at 3.35 TB/s bandwidth. High Bandwidth Memory provides the throughput needed to keep tensor cores fed during large batch inference.

Interconnect

NVLink 4.0 (900 GB/s). Enables multi-GPU tensor parallelism with high-bandwidth, low-latency GPU-to-GPU communication.

Precision Performance

PrecisionTFLOPSUse Case
FP81,979Training & inference with mixed-precision
FP16989Full-precision training, fine-tuning, evaluation

Break-Even ROI Calculator

Monthly hours where reserved pricing beats spot for NVIDIA H100 SXM5. Above the break-even point, reserve commits save money.

Spheron

Spot Rate:$1.89/hr
Reserved Rate:$1.61/hr
Break-Even:320 hrs/mo
Monthly Savings:$204/mo

RunPod

Spot Rate:$4.72/hr
Reserved Rate:$4.01/hr
Break-Even:320 hrs/mo
Monthly Savings:$510/mo

Lambda Labs

Spot Rate:$2.99/hr
Reserved Rate:$2.54/hr
Break-Even:320 hrs/mo
Monthly Savings:$323/mo

Break-even at 320 hours/month: if you run NVIDIA H100 SXM5 more than 320 hours per month, reserved pricing on all three providers saves money. At 720 hours/month (24/7), reserved saves $510+/mo vs on-demand.

Related GPUs

Frequently Asked Questions

How much does it cost to rent an NVIDIA H100 SXM5 per hour?β–Ύ
The lowest spot rate for NVIDIA H100 SXM5 is $1.89/hr. On-demand pricing starts at $4.72/hr depending on the provider and region. Reserved commitments can reduce costs by ~15%.
What is the monthly reserved pricing for NVIDIA H100 SXM5?β–Ύ
Monthly reserved pricing for NVIDIA H100 SXM5 ranges from $2,889/mo to $3,398/mo depending on commitment level and provider.
Can an NVIDIA H100 SXM5 run 70B parameter LLMs?β–Ύ
NVIDIA H100 SXM5 with 80GB HBM3 VRAM can run 70B parameter models with tensor parallelism across 2-4 GPUs. At FP8 precision, the model fits within 80GB HBM3 with room for KV-cache.
Is spot pricing reliable for distributed NVIDIA H100 SXM5 training?β–Ύ
Spot pricing offers 30-60% savings but carries preemption risk. For distributed training, implement SIGTERM handlers with S3 checkpointing every 30 minutes. Vast.ai and RunPod provide 30-second eviction warnings. For critical workloads, reserved pricing eliminates preemption entirely.

Compare alternatives & next steps

Hardware Analysis & Practical Guides

What should I do next?