Llama 3.3 70B on a Single H200 at 128K: FP8 vLLM Config Guide
Step-by-step vLLM launch config to serve Llama 3.3 70B at full 128K context on one NVIDIA H200 SXM5. FP8 weights, FP8 KV-cache flags, and the memory budget that makes (or breaks) single-GPU serving.
Direct answer
Yes โ a single NVIDIA H200 SXM5 (141 GB) can serve Llama 3.3 70B at 128K context, but only if both weights and KV-cache use FP8 precision. A vLLM launch with --quantization fp8 for weights plus --kv-cache-dtype fp8 for the KV-cache lands the total memory budget at approximately 100โ110 GB โ within the 141 GB envelope with headroom. Drop the KV-cache flag and it defaults to FP16, pushing the total to ~135 GB: borderline or over, with no room for overhead or batch size growth.
This article shows the exact flags, the memory math, and the failure mode that breaks most single-H200 deployments.
Why single-H200 serving matters
Llama 3.3 70B is 70.6B parameters. At FP16 (2 bytes/param), weights alone are 140 GB โ over the H100's 80 GB and requiring 2รH100 with tensor parallelism. FP8 (1 byte/param) halves weights to 71 GB, which fits a single H200 and sidesteps NVLink and TP complexity entirely.
The complication is KV-cache precision. For any non-trivial context, the KV-cache competes with weights for VRAM:
total VRAM = model_weights + kv_cache + 1.2 GB CUDA overhead + 0.4 GB activations + 10% fragmentation
If weights are FP8 but KV-cache defaults to FP16, the gains from weight quantization are partially eaten by doubled KV-cache memory.
Memory budget: weights + KV cache + overhead
Using the canonical VRAM engine from Llama 3.3 70B VRAM requirements and vLLM's documented quantization modes:
| Component | Precision | Size @ 128K context |
|---|---|---|
| Model weights (70.6B params) | FP8 | ~71 GB |
| Model weights (70.6B params) | FP16 | ~140 GB |
| KV-cache (batch=1, 128K tokens) | FP8 | 7โ20 GB* |
| KV-cache (batch=1, 128K tokens) | FP16 | 14โ40 GB* |
| CUDA overhead + activations | โ | ~1.6 GB |
| Fragmentation headroom (10%) | โ | 8โ16 GB |
* Layer-count uncertainty: the registry lists 28 layers for Llama 3.3 70B but 80 for Llama 3.1 70B (same architecture). We size 70B KV-cache against the 80-layer figure to avoid under-provisioning: FP8 KV โ 20 GB, FP16 KV โ 40 GB. See KV-cache breakdown for the formula: KV cache (GB) = 2 ร layers ร kv_heads ร head_dim ร context ร bytes รท 1,073,741,824.
Scenario A โ FP8 weights + FP8 KV-cache (recommended):
71 + 20 + 1.6 + (71 + 20 + 1.6) ร 0.10 โ 101 GB โ fits 1ร H200 (141 GB) with ~40 GB headroom.
Scenario B โ FP8 weights + FP16 KV-cache (default, dangerous):
71 + 40 + 1.6 + (71 + 40 + 1.6) ร 0.10 โ 125 GB โ fits, but headroom drops to ~16 GB, and batch>1 or context>128K pushes past 141 GB.
These are theoretical estimates derived from the deterministic VRAM engine. We have not benchmarked this specific config on physical H200 hardware. Validate before production deployment.
The vLLM command
Prerequisites:
- vLLM
0.6.0+(FP8 W8A8 and--kv-cache-dtypesupport; verified against vLLM quantization docs) - Llama 3.3 70B converted to an FP8 checkpoint (via llm-compressor or NVIDIA Model Optimizer)
- H200 SXM5 (141 GB HBM3e, Hopper FP8 support confirmed in vLLM's hardware compatibility table)
vllm serve meta-llama/Llama-3.3-70B-Instruct \
--quantization fp8 \
--kv-cache-dtype fp8 \
--max-model-len 128000 \
--tensor-parallel-size 1 \
--port 8000 \
--api-key YOUR_API_KEY
If you serve an FP8 checkpoint (e.g. an llm-compressor or ModelOpt export), vLLM auto-detects the weight quantization from
config.json;--quantization fp8is only required to dynamically quantize a BF16 checkpoint at load time. Do not also pass--dtype float8โ that forces the activation dtype, which is a separate setting and is not needed for FP8 weight serving.
Flag explanation:
--quantization fp8โ load FP8-quantized weights (W8A8)--kv-cache-dtype fp8โ critical: quantize KV-cache to FP8 instead of defaulting to the model's native FP16--max-model-len 128000โ native context is 131,072 (128K); 128000 leaves a small margin--tensor-parallel-size 1โ single GPU (no TP split)
To check your config's memory needs against the VRAM calculator: Calculator for Llama 3.3 70B @ 128K FP8.
The KV-cache gotcha
vLLM's default kv_cache_dtype is "auto" โ which resolves to the model's native precision (FP16 for Llama 3.3). Many deployments set --quantization fp8 for weights and assume KV-cache inherits the savings. It does not โ without --kv-cache-dtype fp8, KV-cache stays FP16 and the memory budget reverts to Scenario B.
This is why a configuration that should fit on one H200 can fail at 8โ10% GPU memory utilization: the weights are small, but the FP16 KV-cache at 128K context is larger than expected.
Validation steps
- Before launch โ confirm the checkpoint declares FP8 quantization in its config:
python -c "import json; qc = json.load(open('config.json')).get('quantization_config'); print(qc)" # Expect a dict like {'quant_method': 'fp8', ...}, not None - Launch and monitor VRAM:
vllm serve meta-llama/Llama-3.3-70B-Instruct --quantization fp8 --kv-cache-dtype fp8 ... nvidia-smi --query-gpu=memory.used,memory.total --format=csv -l 1 - Run a 128K-context test โ feed 128K tokens and check utilization stays below 110 GB (Scenario A budget):
curl -X POST http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer YOUR_API_KEY" \ -d '{"model":"meta-llama/Llama-3.3-70B-Instruct","messages":[{"role":"user","content":"<paste 128K tokens of context>"}],"max_tokens":1}' - Verify KV-cache dtype in vLLM logs โ search for
KV cache dtype: fp8.
Concurrency and limitations
- Batch=1 headroom: at FP8 weights + FP8 KV-cache, 1รH200 uses ~101 GB of 141 GB โ roughly 40 GB free, enough for about 2 additional 128K-context requests (โ20 GB FP8 KV each) before you approach capacity. Beyond that, scale to 2รH200 with
tp_size=2. - Hopper-only: FP8 W8A8 quantization requires Hopper (H100/H200/B200) or Ada (RTX 4090). On older architectures, fall back to INT4 โ see quantization formats guide.
- Quality trade-off: FP8 W8A8 typically lands within ~1% of BF16 on standard LLM benchmarks (community and vendor reports), but quantized KV-cache can degrade long-context coherence. Validate output quality on your representative prompts โ we do not publish a single universal perplexity delta.
- MLA models: DeepSeek R1 compresses its KV-cache via Multi-Latent Attention (~4 GB at 128K context โ small), but its 671B FP8 weights alone are ~671 GB, so single-H200 serving of R1 is not possible at full weight residency; production R1 needs multi-GPU (4รH200 INT4 / 6รH200 FP8). See the DeepSeek R1 serving guide and FP8 vs INT4 memory analysis.
Common failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| OOM at ~110 GB used | KV-cache defaulted to FP16 | Add --kv-cache-dtype fp8 |
| "Quantization method fp8 not supported" | vLLM < 0.6.0 or non-Hopper GPU | Upgrade vLLM or use --quantization bitsandbytes with W4A16 |
| Slow first-token at high context | GPU page faults / HBM initialization | Warm up with a short prompt before the 128K request |
| Accuracy degradation >0.5% | Aggressive FP8 calibration | Use llm-compressor with dataset calibration instead of no-calibration defaults |
Related resources
- Llama 3.3 70B VRAM requirements โ model-specific VRAM table
- The KV-cache memory trap โ why context length scales costly
- NVIDIA H200 SXM5: 141GB HBM3e for large-context inference โ hardware specs and throughput
- LLM VRAM Calculator โ compute your exact config
- Quantization formats: FP8 vs INT4 vs AWQ โ precision trade-offs
- vLLM production deployment guide โ multi-GPU scaling beyond this config
Related Articles
Mixture of Experts (MoE) Explained: Memory and Serving Math
What Mixture of Experts (MoE) means for GPU memory: why 671B total but 37B active still needs all weights resident, plus bandwidth math.
ArchitectureB300 Blackwell Ultra: Full Architecture Deep Dive and VRAM Analysis
B300 Blackwell Ultra: 288GB HBM3e, 9 TB/s, 2,500 FP8 TFLOPS โ NVIDIA's target spec per the Blackwell Ultra whitepaper. VRAM analysis: what fits on 3x B300 vs 5x B200, and why pre-production means verify before buying.
ArchitectureDeepSeek R1 671B: FP8 vs INT4 Memory Overhead & PPL Degradation
DeepSeek R1 671B: 671 GB weights (FP8) vs 336 GB (INT4). MLA KV cache of 4.19 GB at 128K. H200/B200 cluster sizing at $8.37/hr (INT4) vs $16.74/hr (FP8).