NVIDIA H100 SXM5 Cloud Pricing & Specs (2026)
The H100 SXM5 is the workhorse of AI training clusters. Its 80 GB HBM3 memory handles 70B parameter models with tensor parallelism across 2-4 GPUs, while 3.35 TB/s bandwidth keeps the compute units fed during all-reduce operations.
Target workload: Multi-node 70B+ LLM pre-training & fine-tuning
80 GB VRAM
Manufacturer specification for onboard memory.
3.35 TB/s
Peak memory bandwidth from the manufacturer specification.
NVIDIA H100 SXM5: Key Numbers at a Glance
Rental cost: no verified rate rows are tracked yet — this page asserts specifications only, never a price.
70B fit: Fits Llama 3.3 70B at FP8/INT8 (weights ≈71 GB at 1 byte/param) with a short-context KV budget; FP16 (140 GB) requires 2-way tensor parallelism.
Bottleneck: Memory-bandwidth bound at large batch; compute-bound at small batch with FP8/BF16
OpenGPU Radar does not currently track "H100 SXM5" rows in data/providers.json. No price is asserted on this page — the specifications below are manufacturer-sourced.
Compare GPUs with tracked rates →Compatible Models for NVIDIA H100 SXM5
Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 80 GB. FP16 shows where full precision also fits single-GPU.
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
Specifications
H100 SXM5 Memory-Bandwidth Ceiling Analysis: In an 8-way HGX baseboard, each GPU connects via NVLink 4.0 at 900 GB/s bidirectional bandwidth. For Llama 70B FP8 tensor parallelism, TP=2 splits the 72 GB model across 2 GPUs (36 GB each) with 16 GB headroom for KV-cache at 32k context. TP=4 splits across 4 GPUs (18 GB each) but introduces 3 all-reduce sync points per forward pass — the 900 GB/s NVLink mesh absorbs this with <5% communication overhead. At TP=8, the model shards to 9 GB per GPU, but NCCL all-reduce latency increases to ~8% due to the ring-all-reduce topology across 8 nodes. The optimal config is TP=2 for inference (lowest latency) and TP=4 for training (better compute utilization). At batch≥128, the H100's 3.35 TB/s HBM3 bandwidth becomes the limiting factor: the GPU's 1,979 FP8 TFLOPS of compute throughput can only be sustained if the memory subsystem delivers sufficient weight data per cycle. With 80 GB HBM3 at 3.35 TB/s, the effective bandwidth-to-compute ratio is ~1.69 bytes per FLOP, well above the theoretical minimum of ~0.5 bytes/FLOP for FP8 inference, confirming the memory-bandwidth bound.
Next steps
Compare Alternatives
Head-to-head comparisons against this GPU — specs, observed hourly rates, and the workload verdict.
The H200 eliminates 2-GPU tensor sharding for Llama 70B at 128k context. The H100 remains the lower-cost option for sub-32k workloads at $1.89/hr vs $2.50/hr.
The B200 delivers 2x inference throughput via FP4 but requires liquid-cooled infrastructure at 1000W TDP. The H100 is deployment-ready in air-cooled datacenters today.
H100 SXM5 leads FP8-native serving: 1,979 FP8 TFLOPS and 3.35 TB/s bandwidth deliver a higher decode ceiling per GPU, and NVLink 4.0 at 900 GB/s scales tensor parallelism. The A100 80GB is 16% cheaper per hour observed ($1.59 vs $1.89) and remains the budget choice for INT8/INT4 batch workloads and Ampere-standardized tooling — but it has no FP8 tensor cores, so Hopper's native FP8 pipeline advantage does not exist on Ampere.
H100 SXM5 is the cost-efficient workhorse: identical 1,979 FP8 TFLOPS and NVLink 4.0 at $1.89/hr observed — $0.90/hr less than the H200. The H200's 141 GB (+76% VRAM, +43% bandwidth) is the unlock for long-context work: 70B at FP8 with full 128k KV-cache fits one GPU, where the H100's 80 GB forces eviction or shorter context. Choose H100 for cost-efficient standard serving; choose H200 when context length or single-GPU simplicity outruns.
B200 doubles the platform: 192 GB fits 70B at BF16-class precision where the H100 needs FP8, and 8.0 TB/s plus 2,250 FP8 TFLOPS push per-GPU throughput past Hopper. But it costs 2.1× per hour observed ($3.99 vs $1.89). Choose H100 unless you need single-GPU BF16 70B, FP4-capable kernels, or cluster consolidation — the B200's per-GPU advantage only pays when its throughput is actually sustained.