Spec & Pricing Reference

NVIDIA B200 Blackwell Cloud Pricing & Specs (2026)

The B200 Blackwell delivers 192 GB HBM3e at 8 TB/s with NVLink 5.0 at 1.8 TB/s. Its 4,500 FP4 TFLOPS enable next-generation MoE training and extreme-throughput inference for trillion-parameter models.

Target workload: Next-gen MoE training & extreme-throughput FP4 serving

Memory192GB HBM3e
Bandwidth8.0 TB/s
FP8 TFLOPS2,250
Observed rate$0.00/hr
192 GB VRAM
Manufacturer SpecHIGH
SourceNVIDIA manufacturer specifications
VerifiedSep 30, 2026
Value
192 GB
Methodology

Manufacturer specification for onboard memory.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
8.0 TB/s
Manufacturer SpecHIGH
SourceNVIDIA specification
VerifiedSep 30, 2026
Value
8.0 TB/s GB/s
Methodology

Peak memory bandwidth from the manufacturer specification.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
$0.00/hr on-demand
Observed On-Demand RateHIGH
SourceObserved provider API rate (lowest on-demand row)
VerifiedSep 30, 2026
Value
0 USD/hr
Methodology

Lowest on-demand hourly row for "B200" across tracked providers in data/providers.json; refreshed daily.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
Methodology โ†’

NVIDIA B200 Blackwell: Key Numbers at a Glance

Rental cost: lowest observed on-demand $0.00/hr, spot rows from $0.00/hr across 4 tracked providers (refreshed daily).

70B fit: Fits Llama 3.3 70B at FP16 (140 GB min VRAM) with KV-cache headroom for 128k context.

Bottleneck: Memory-bandwidth bound at large batch; compute-bound at small batch with FP8/BF16

Observed Pricing for NVIDIA B200 Blackwell

Spot and on-demand rows for B200 refreshed from provider APIs. Lowest on-demand: $0.00/hr.

ProviderGPU & VRAMInterconnectSpot Price ($/hr)On-Demand ($/hr)Monthly ($/720h)StatusAction
Lambda Labs
N/A$0/hr
Calculated EstimateMEDIUM
SourceLambda Labs
VerifiedSep 26, 2026
Value
0 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5ร— spot (conservative; actual range 1.8ร—โ€“3.2ร—).

Assumptions & Parameters
  • range: 1.8ร—โ€“3.2ร—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$0 / hr$0 / moInstant
Vast.ai
N/A$4.2/hr
Calculated EstimateMEDIUM
SourceVast.ai
VerifiedSep 26, 2026
Value
4.2 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5ร— spot (conservative; actual range 1.8ร—โ€“3.2ร—).

Assumptions & Parameters
  • range: 1.8ร—โ€“3.2ร—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$5.8 / hr$0 / moInstant
Spheron
N/A$5.37/hr
Calculated EstimateMEDIUM
SourceSpheron
VerifiedSep 26, 2026
Value
5.37 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5ร— spot (conservative; actual range 1.8ร—โ€“3.2ร—).

Assumptions & Parameters
  • range: 1.8ร—โ€“3.2ร—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$7.2 / hr$0 / moInstant
RunPod
N/A$6.79/hr
Calculated EstimateMEDIUM
SourceRunPod
VerifiedSep 26, 2026
Value
6.79 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5ร— spot (conservative; actual range 1.8ร—โ€“3.2ร—).

Assumptions & Parameters
  • range: 1.8ร—โ€“3.2ร—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$6.79 / hr$0 / moInstant
Monthly estimates use continuous spot rate ร— 720h. On-demand calculated at 2.5ร— spot (range: 1.8ร—โ€“3.2ร—). All rates subject to preemption and provider availability.
Data Freshness: Public Cloud APIs & Market Scraping | Refreshed Daily (UTC)Benchmark Baseline: Ubuntu 24.04, CUDA 12.4, vLLM v0.6.x, PagedAttention v2, FlashAttention-3

Compatible Models for NVIDIA B200 Blackwell

Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 192 GB. FP16 shows where full precision also fits single-GPU.

Browse all model VRAM pages โ†’

Specifications

ArchitectureBlackwell GB200 โ€” 4NP TSMC
Memory192GB HBM3e
Bandwidth8.0 TB/s
InterconnectNVLink 5.0 (1.8 TB/s)
TDP1000W
FP16 TFLOPS1,125
FP8 TFLOPS2,250
FP4 TFLOPS4,500
Recommended quantizationFP4 / FP8 native โ€” FP4 cuts memory 50% with <5% quality loss
Best cluster topology8-way NVLink 5.0 Full Mesh (1.8 TB/s per GPU)

B200 Blackwell FP4 Precision Scaling & Liquid Cooling Analysis: The B200's second-generation Transformer Engine adds native FP4 precision โ€” 4,500 TFLOPS at one-quarter the precision of FP16. FP4 quantization cuts memory requirements in half: Llama 70B at FP4 occupies ~35 GB (vs 70 GB FP16), fitting entirely on a single B200 with 157 GB remaining for KV-cache. The quality tradeoff is <5% perplexity degradation on standard benchmarks, acceptable for inference workloads. NVLink 5.0 at 1.8 TB/s per GPU enables 8-way tensor parallelism with near-zero communication overhead โ€” the doubled bandwidth vs NVLink 4.0 absorbs all-reduce latency even at batch=256. The critical infrastructure dependency: B200 at 1000W TDP requires liquid-cooled racks. Air-cooled datacenters cannot sustain the thermal envelope. Providers offering B200 must have direct-to-chip liquid cooling infrastructure, which limits availability to purpose-built AI datacenters.

Next steps

Compare Alternatives

Head-to-head comparisons against this GPU โ€” specs, observed hourly rates, and the workload verdict.

Related GPUs

Frequently Asked Questions

How much does it cost to rent NVIDIA B200 Blackwell per hour?โ–พ
The lowest observed on-demand rate for NVIDIA B200 Blackwell is $0.00/hr (4 tracked providers); observed spot rows start at $0.00/hr. Rates refresh daily from provider APIs.
Can NVIDIA B200 Blackwell run 70B parameter LLMs?โ–พ
Fits Llama 3.3 70B at FP16 (140 GB min VRAM) with KV-cache headroom for 128k context.
What is NVIDIA B200 Blackwell's memory bandwidth?โ–พ
NVIDIA B200 Blackwell has 8.0 TB/s of memory bandwidth across 192GB HBM3e. Token generation is memory-bandwidth bound at batch size 1, so decode throughput scales with this figure โ€” no tokens/sec value is claimed without a benchmark row.
Rates observed across 4 providers; refreshed daily (UTC).Spec sources: NVIDIA manufacturer specifications.Methodology โ†’