NVIDIA B200 Blackwell Cloud Pricing & Specs (2026)
The B200 Blackwell delivers 192 GB HBM3e at 8 TB/s with NVLink 5.0 at 1.8 TB/s. Its 4,500 FP4 TFLOPS enable next-generation MoE training and extreme-throughput inference for trillion-parameter models.
Target workload: Next-gen MoE training & extreme-throughput FP4 serving
192 GB VRAM
Manufacturer specification for onboard memory.
8.0 TB/s
Peak memory bandwidth from the manufacturer specification.
$0.00/hr on-demand
Lowest on-demand hourly row for "B200" across tracked providers in data/providers.json; refreshed daily.
NVIDIA B200 Blackwell: Key Numbers at a Glance
Rental cost: lowest observed on-demand $0.00/hr, spot rows from $0.00/hr across 4 tracked providers (refreshed daily).
70B fit: Fits Llama 3.3 70B at FP16 (140 GB min VRAM) with KV-cache headroom for 128k context.
Bottleneck: Memory-bandwidth bound at large batch; compute-bound at small batch with FP8/BF16
Observed Pricing for NVIDIA B200 Blackwell
Spot and on-demand rows for B200 refreshed from provider APIs. Lowest on-demand: $0.00/hr.
Featured GPU Pricing Pages
80GB HBM3 โข 3.35 TB/s โข NVLink 4.0
40GB HBM2 โข 1.55 TB/s โข NVLink 3.0
Verified spot rates & reserved pricing
24GB GDDR6 โข 0.62 TB/s โข PCIe 4.0
24GB GDDR6X โข 0.94 TB/s โข PCIe 4.0
16GB GDDR6 โข 0.32 TB/s โข PCIe 3.0
Compatible Models for NVIDIA B200 Blackwell
Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 192 GB. FP16 shows where full precision also fits single-GPU.
VRAM breakdown โ
VRAM breakdown โ
VRAM breakdown โ
VRAM breakdown โ
VRAM breakdown โ
VRAM breakdown โ
VRAM breakdown โ
VRAM breakdown โ
Specifications
B200 Blackwell FP4 Precision Scaling & Liquid Cooling Analysis: The B200's second-generation Transformer Engine adds native FP4 precision โ 4,500 TFLOPS at one-quarter the precision of FP16. FP4 quantization cuts memory requirements in half: Llama 70B at FP4 occupies ~35 GB (vs 70 GB FP16), fitting entirely on a single B200 with 157 GB remaining for KV-cache. The quality tradeoff is <5% perplexity degradation on standard benchmarks, acceptable for inference workloads. NVLink 5.0 at 1.8 TB/s per GPU enables 8-way tensor parallelism with near-zero communication overhead โ the doubled bandwidth vs NVLink 4.0 absorbs all-reduce latency even at batch=256. The critical infrastructure dependency: B200 at 1000W TDP requires liquid-cooled racks. Air-cooled datacenters cannot sustain the thermal envelope. Providers offering B200 must have direct-to-chip liquid cooling infrastructure, which limits availability to purpose-built AI datacenters.
Next steps
Compare Alternatives
Head-to-head comparisons against this GPU โ specs, observed hourly rates, and the workload verdict.
The B200 delivers 2x inference throughput via FP4 but requires liquid-cooled infrastructure at 1000W TDP. The H100 is deployment-ready in air-cooled datacenters today.
B200 doubles the platform: 192 GB fits 70B at BF16-class precision where the H100 needs FP8, and 8.0 TB/s plus 2,250 FP8 TFLOPS push per-GPU throughput past Hopper. But it costs 2.1ร per hour observed ($3.99 vs $1.89). Choose H100 unless you need single-GPU BF16 70B, FP4-capable kernels, or cluster consolidation โ the B200's per-GPU advantage only pays when its throughput is actually sustained.