Architecture2026-10-05•By Sree•5 min read

B300 Blackwell Ultra: Full Architecture Deep Dive and VRAM Analysis

B300 Blackwell Ultra: 288GB HBM3e, 9 TB/s, 2,500 FP8 TFLOPS — NVIDIA's target spec per the Blackwell Ultra whitepaper. VRAM analysis: what fits on 3x B300 vs 5x B200, and why pre-production means verify before buying.

Status: B300 specifications below come from the NVIDIA Blackwell Ultra Architecture whitepaper — manufacturer target spec, pre-production (the site's provenance record for b300-blackwell-ultra-rates). The GPU is not yet generally shipping, and no B300 cloud rate exists in providers.json as of 2026-10-06. Every number is labeled accordingly.

Direct answer

B300 Blackwell Ultra targets 288 GB HBM3e, 9.0 TB/s memory bandwidth, 2,500 FP8 TFLOPS, and 1,200W TDP — a 50% VRAM increase over B200's 192 GB, not the doubling earlier projections assumed. In practical terms:

  • DeepSeek R1 671B FP8 (786 GB @128K context): 3× B300 (864 GB) vs 5× B200 (960 GB) vs 6× H200 (846 GB)
  • Llama 3.3 70B FP8 (81 GB @32K): single B300 or single B200 or single H200 — B300's extra VRAM does not change this class
  • The headline claim from older write-ups — "384 GB, single-GPU for 700B+ models" — is not supported by the whitepaper spec. At 288 GB, R1 FP8 still needs 3 GPUs.

For current buying decisions, use the verified B200 pricing and the minimum viable cluster analysis for DeepSeek R1.

Key specifications

MetricB200 SXM (current)B300 Ultra (pre-production target)Status
VRAM192 GB HBM3e288 GB HBM3ewhitepaper target spec
Memory bandwidth8.0 TB/s9.0 TB/swhitepaper target spec
FP8 compute2,250 TFLOPS2,500 TFLOPSwhitepaper target spec
FP4 compute4,500 TFLOPS5,000 TFLOPSwhitepaper target spec
TDP1,000W1,200Wwhitepaper target spec
Site throughput record (70B FP8, batch=1)~180 tok/s~200 tok/ssite GPU spec record
Cloud price$3.99/hr Vast.ai (2026-10-03)no rate publishedabsent from providers.json

Architecture evolution: B200 → B300

B200 SXM (available now, observed in providers.json):

  • 192 GB HBM3e, 8.0 TB/s
  • NVLink 5.0 at 1.8 TB/s per GPU
  • 2,250 FP8 TFLOPS, 1,000W — requires liquid cooling
  • 8× B200 node: 1,536 GB, $31.92/hr (Vast.ai)

B300 Blackwell Ultra (whitepaper target spec, pre-production):

  • 288 GB HBM3e (+50% capacity), 9.0 TB/s (+12.5% bandwidth)
  • 2,500 FP8 TFLOPS (+11%) and 5,000 FP4 TFLOPS
  • 1,200W TDP — same liquid-cooling requirement, higher thermal budget
  • Positioned for frontier training and single-socket capacity records — but 288 GB is still less than half of R1's FP8 weights-plus-KV footprint

VRAM analysis: What models fit on B300?

All model figures from the canonical VRAM engine (weights + KV + overhead, 10% fragmentation), context stated where it matters:

ModelFP8 VRAM (context)Fits on B300 (288 GB)?Fits on B200 (192 GB)?
Llama 3.3 70B81 GB @32K1 GPU1 GPU
Qwen 2.5 72B~87 GB @32K1 GPU1 GPU
Llama 3.1 405B482 GB @32K · 586 GB @128K2 GPU @32K · 3 GPU @128K3 GPU @32K · 4 GPU @128K
DeepSeek R1 671B786 GB @128K3 GPU (864 GB)5 GPU (960 GB)
DeepSeek V3 687B~807 GB @128K (derived, 687B params, same engine)3 GPU (864 GB)5 GPU (960 GB)

Older revisions of this table listed 70B at "140 GB" and R1 at "1,340 GB" — both were FP16 weights mislabeled as FP8 service totals. FP8 service totals include KV cache and overhead: 81 GB and 786 GB respectively (see VRAM calculator).

The capacity story: 288 GB lets a B300 hold any 70B-class model with massive KV headroom, and cuts R1 FP8 from 5 GPUs (B200) to 3 — a 40% reduction in GPU count for the flagship workload. It does not enable single-GPU 700B-class FP8 inference; that claim required the unsupported 384 GB projection.

Performance projections

WorkloadB200 SXM recordB300 targetBasis
Llama 70B FP8, batch=1~180 tok/s per GPU~200 tok/s per GPUsite GPU spec records (+11% FP8 TFLOPS)
8-GPU node, 70B replicas1,440 tok/s (8×180)~1,600 tok/s (8×200) if 8-GPU config existsderived
DeepSeek R1 671Bno site throughput recordno record—

Throughput beyond the per-GPU records is derived from spec ratios, not measured — B300 silicon is pre-production. Ignore any pre-launch write-up quoting "verified" multi-thousand-tok/s R1 figures; no such benchmark exists in the project's data.

Deployment scenarios

3x B300 (hypothetical — pre-production, no published price)

  • Total VRAM: 864 GB vs R1 FP8's 786 GB needed (78 GB spare — 4× is the comfortable config at 1,152 GB)
  • Caveat: pricing unknown; 4×B300 could plausibly cost more than 5×B200 ($19.95/hr) until rates appear in providers.json

5x B200 (verified single-node alternative — $19.95/hr)

  • Total VRAM: 960 GB (174 GB spare over the 786 GB needed)
  • Hourly cost: $3.99 × 5 = $19.95/hr (Vast.ai, observed)
  • Alternative: 6× H200 = 846 GB at $16.74/hr — still the price-performance pick (see minimum viable cluster)

8x B200 (single-node baseline)

  • Total VRAM: 1,536 GB — 750 GB spare over the 786 GB needed; everything fits one node
  • Availability: shipping today (B300 not yet available)

What changes the result?

  • B300 pricing: until a rate lands in providers.json, cost-per-token math for B300 is speculation — do not model it
  • Actual vs target spec: pre-production whitepaper numbers have moved before (Hopper, Blackwell proper) — treat 288 GB / 2,500 TFLOPS as best-known, not final
  • Context length: R1's 786 GB assumes 128K; at ≤4K it's 741 GB — still 3× B300, still 5× B200 (4× B200 = 768 GB only fits short-context R1, marginally)
  • Software readiness: NVLink/NCCL support for new silicon typically lags first hardware availability

Alternatives — all verified data

  • B200 SXM now: Spheron ($4.49/hr) or Vast.ai ($3.99/hr) (providers.json, 2026-10-03)
  • H100 SXM5: $1.89/hr Vast.ai — single GPU serves models ≤80 GB FP8 (TP=2 needed for 70B FP8's 81 GB)
  • H200: $2.79/hr Vast.ai — 141 GB, single-GPU 70B FP8, the pragmatic bridge until B300 rates exist

Conclusion

Decision factorRecommendation
Need R1 FP8 inference now6× H200 at $16.74/hr (cheapest verified) or 5× B200 at $19.95/hr
Need 70B FP8, single GPUH200 ($2.79/hr) or B200 ($3.99/hr) — B300 unnecessary
Planning next-gen capacityWatch for B300 rates in providers.json before modeling anything
Budget-constrainedH100 at $1.89/hr with TP=2 for 70B FP8, or INT4 (42.5 GB) single-GPU
Risk toleranceB300 is pre-production target spec — production workloads on verified B200/H200 today

Calculate VRAM requirements for your model with the GPU VRAM Calculator. Current B200 rates are live at NVIDIA B200 GPU specs.