B300 Blackwell Ultra: Full Architecture Deep Dive and VRAM Analysis
B300 Blackwell Ultra: 288GB HBM3e, 9 TB/s, 2,500 FP8 TFLOPS — NVIDIA's target spec per the Blackwell Ultra whitepaper. VRAM analysis: what fits on 3x B300 vs 5x B200, and why pre-production means verify before buying.
Status: B300 specifications below come from the NVIDIA Blackwell Ultra Architecture whitepaper — manufacturer target spec, pre-production (the site's provenance record for
b300-blackwell-ultra-rates). The GPU is not yet generally shipping, and no B300 cloud rate exists inproviders.jsonas of 2026-10-06. Every number is labeled accordingly.
Direct answer
B300 Blackwell Ultra targets 288 GB HBM3e, 9.0 TB/s memory bandwidth, 2,500 FP8 TFLOPS, and 1,200W TDP — a 50% VRAM increase over B200's 192 GB, not the doubling earlier projections assumed. In practical terms:
- DeepSeek R1 671B FP8 (786 GB @128K context): 3× B300 (864 GB) vs 5× B200 (960 GB) vs 6× H200 (846 GB)
- Llama 3.3 70B FP8 (81 GB @32K): single B300 or single B200 or single H200 — B300's extra VRAM does not change this class
- The headline claim from older write-ups — "384 GB, single-GPU for 700B+ models" — is not supported by the whitepaper spec. At 288 GB, R1 FP8 still needs 3 GPUs.
For current buying decisions, use the verified B200 pricing and the minimum viable cluster analysis for DeepSeek R1.
Key specifications
| Metric | B200 SXM (current) | B300 Ultra (pre-production target) | Status |
|---|---|---|---|
| VRAM | 192 GB HBM3e | 288 GB HBM3e | whitepaper target spec |
| Memory bandwidth | 8.0 TB/s | 9.0 TB/s | whitepaper target spec |
| FP8 compute | 2,250 TFLOPS | 2,500 TFLOPS | whitepaper target spec |
| FP4 compute | 4,500 TFLOPS | 5,000 TFLOPS | whitepaper target spec |
| TDP | 1,000W | 1,200W | whitepaper target spec |
| Site throughput record (70B FP8, batch=1) | ~180 tok/s | ~200 tok/s | site GPU spec record |
| Cloud price | $3.99/hr Vast.ai (2026-10-03) | no rate published | absent from providers.json |
Architecture evolution: B200 → B300
B200 SXM (available now, observed in providers.json):
- 192 GB HBM3e, 8.0 TB/s
- NVLink 5.0 at 1.8 TB/s per GPU
- 2,250 FP8 TFLOPS, 1,000W — requires liquid cooling
- 8× B200 node: 1,536 GB, $31.92/hr (Vast.ai)
B300 Blackwell Ultra (whitepaper target spec, pre-production):
- 288 GB HBM3e (+50% capacity), 9.0 TB/s (+12.5% bandwidth)
- 2,500 FP8 TFLOPS (+11%) and 5,000 FP4 TFLOPS
- 1,200W TDP — same liquid-cooling requirement, higher thermal budget
- Positioned for frontier training and single-socket capacity records — but 288 GB is still less than half of R1's FP8 weights-plus-KV footprint
VRAM analysis: What models fit on B300?
All model figures from the canonical VRAM engine (weights + KV + overhead, 10% fragmentation), context stated where it matters:
| Model | FP8 VRAM (context) | Fits on B300 (288 GB)? | Fits on B200 (192 GB)? |
|---|---|---|---|
| Llama 3.3 70B | 81 GB @32K | 1 GPU | 1 GPU |
| Qwen 2.5 72B | ~87 GB @32K | 1 GPU | 1 GPU |
| Llama 3.1 405B | 482 GB @32K · 586 GB @128K | 2 GPU @32K · 3 GPU @128K | 3 GPU @32K · 4 GPU @128K |
| DeepSeek R1 671B | 786 GB @128K | 3 GPU (864 GB) | 5 GPU (960 GB) |
| DeepSeek V3 687B | ~807 GB @128K (derived, 687B params, same engine) | 3 GPU (864 GB) | 5 GPU (960 GB) |
Older revisions of this table listed 70B at "140 GB" and R1 at "1,340 GB" — both were FP16 weights mislabeled as FP8 service totals. FP8 service totals include KV cache and overhead: 81 GB and 786 GB respectively (see VRAM calculator).
The capacity story: 288 GB lets a B300 hold any 70B-class model with massive KV headroom, and cuts R1 FP8 from 5 GPUs (B200) to 3 — a 40% reduction in GPU count for the flagship workload. It does not enable single-GPU 700B-class FP8 inference; that claim required the unsupported 384 GB projection.
Performance projections
| Workload | B200 SXM record | B300 target | Basis |
|---|---|---|---|
| Llama 70B FP8, batch=1 | ~180 tok/s per GPU | ~200 tok/s per GPU | site GPU spec records (+11% FP8 TFLOPS) |
| 8-GPU node, 70B replicas | 1,440 tok/s (8×180) | ~1,600 tok/s (8×200) if 8-GPU config exists | derived |
| DeepSeek R1 671B | no site throughput record | no record | — |
Throughput beyond the per-GPU records is derived from spec ratios, not measured — B300 silicon is pre-production. Ignore any pre-launch write-up quoting "verified" multi-thousand-tok/s R1 figures; no such benchmark exists in the project's data.
Deployment scenarios
3x B300 (hypothetical — pre-production, no published price)
- Total VRAM: 864 GB vs R1 FP8's 786 GB needed (78 GB spare — 4× is the comfortable config at 1,152 GB)
- Caveat: pricing unknown; 4×B300 could plausibly cost more than 5×B200 ($19.95/hr) until rates appear in
providers.json
5x B200 (verified single-node alternative — $19.95/hr)
- Total VRAM: 960 GB (174 GB spare over the 786 GB needed)
- Hourly cost: $3.99 × 5 = $19.95/hr (Vast.ai, observed)
- Alternative: 6× H200 = 846 GB at $16.74/hr — still the price-performance pick (see minimum viable cluster)
8x B200 (single-node baseline)
- Total VRAM: 1,536 GB — 750 GB spare over the 786 GB needed; everything fits one node
- Availability: shipping today (B300 not yet available)
What changes the result?
- B300 pricing: until a rate lands in
providers.json, cost-per-token math for B300 is speculation — do not model it - Actual vs target spec: pre-production whitepaper numbers have moved before (Hopper, Blackwell proper) — treat 288 GB / 2,500 TFLOPS as best-known, not final
- Context length: R1's 786 GB assumes 128K; at ≤4K it's 741 GB — still 3× B300, still 5× B200 (4× B200 = 768 GB only fits short-context R1, marginally)
- Software readiness: NVLink/NCCL support for new silicon typically lags first hardware availability
Alternatives — all verified data
- B200 SXM now: Spheron ($4.49/hr) or Vast.ai ($3.99/hr) (
providers.json, 2026-10-03) - H100 SXM5: $1.89/hr Vast.ai — single GPU serves models ≤80 GB FP8 (TP=2 needed for 70B FP8's 81 GB)
- H200: $2.79/hr Vast.ai — 141 GB, single-GPU 70B FP8, the pragmatic bridge until B300 rates exist
Conclusion
| Decision factor | Recommendation |
|---|---|
| Need R1 FP8 inference now | 6× H200 at $16.74/hr (cheapest verified) or 5× B200 at $19.95/hr |
| Need 70B FP8, single GPU | H200 ($2.79/hr) or B200 ($3.99/hr) — B300 unnecessary |
| Planning next-gen capacity | Watch for B300 rates in providers.json before modeling anything |
| Budget-constrained | H100 at $1.89/hr with TP=2 for 70B FP8, or INT4 (42.5 GB) single-GPU |
| Risk tolerance | B300 is pre-production target spec — production workloads on verified B200/H200 today |
Calculate VRAM requirements for your model with the GPU VRAM Calculator. Current B200 rates are live at NVIDIA B200 GPU specs.