Architecture2026-10-10โ€ขBy Sreeโ€ข5 min read

Llama 3.3 70B on a Single H200 at 128K: FP8 vLLM Config Guide

Step-by-step vLLM launch config to serve Llama 3.3 70B at full 128K context on one NVIDIA H200 SXM5. FP8 weights, FP8 KV-cache flags, and the memory budget that makes (or breaks) single-GPU serving.

Direct answer

Yes โ€” a single NVIDIA H200 SXM5 (141 GB) can serve Llama 3.3 70B at 128K context, but only if both weights and KV-cache use FP8 precision. A vLLM launch with --quantization fp8 for weights plus --kv-cache-dtype fp8 for the KV-cache lands the total memory budget at approximately 100โ€“110 GB โ€” within the 141 GB envelope with headroom. Drop the KV-cache flag and it defaults to FP16, pushing the total to ~135 GB: borderline or over, with no room for overhead or batch size growth.

This article shows the exact flags, the memory math, and the failure mode that breaks most single-H200 deployments.

Why single-H200 serving matters

Llama 3.3 70B is 70.6B parameters. At FP16 (2 bytes/param), weights alone are 140 GB โ€” over the H100's 80 GB and requiring 2ร—H100 with tensor parallelism. FP8 (1 byte/param) halves weights to 71 GB, which fits a single H200 and sidesteps NVLink and TP complexity entirely.

The complication is KV-cache precision. For any non-trivial context, the KV-cache competes with weights for VRAM:

total VRAM = model_weights + kv_cache + 1.2 GB CUDA overhead + 0.4 GB activations + 10% fragmentation

If weights are FP8 but KV-cache defaults to FP16, the gains from weight quantization are partially eaten by doubled KV-cache memory.

Memory budget: weights + KV cache + overhead

Using the canonical VRAM engine from Llama 3.3 70B VRAM requirements and vLLM's documented quantization modes:

ComponentPrecisionSize @ 128K context
Model weights (70.6B params)FP8~71 GB
Model weights (70.6B params)FP16~140 GB
KV-cache (batch=1, 128K tokens)FP87โ€“20 GB*
KV-cache (batch=1, 128K tokens)FP1614โ€“40 GB*
CUDA overhead + activationsโ€”~1.6 GB
Fragmentation headroom (10%)โ€”8โ€“16 GB

* Layer-count uncertainty: the registry lists 28 layers for Llama 3.3 70B but 80 for Llama 3.1 70B (same architecture). We size 70B KV-cache against the 80-layer figure to avoid under-provisioning: FP8 KV โ‰ˆ 20 GB, FP16 KV โ‰ˆ 40 GB. See KV-cache breakdown for the formula: KV cache (GB) = 2 ร— layers ร— kv_heads ร— head_dim ร— context ร— bytes รท 1,073,741,824.

Scenario A โ€” FP8 weights + FP8 KV-cache (recommended): 71 + 20 + 1.6 + (71 + 20 + 1.6) ร— 0.10 โ‰ˆ 101 GB โ€” fits 1ร— H200 (141 GB) with ~40 GB headroom.

Scenario B โ€” FP8 weights + FP16 KV-cache (default, dangerous): 71 + 40 + 1.6 + (71 + 40 + 1.6) ร— 0.10 โ‰ˆ 125 GB โ€” fits, but headroom drops to ~16 GB, and batch>1 or context>128K pushes past 141 GB.

These are theoretical estimates derived from the deterministic VRAM engine. We have not benchmarked this specific config on physical H200 hardware. Validate before production deployment.

The vLLM command

Prerequisites:

vllm serve meta-llama/Llama-3.3-70B-Instruct \
  --quantization fp8 \
  --kv-cache-dtype fp8 \
  --max-model-len 128000 \
  --tensor-parallel-size 1 \
  --port 8000 \
  --api-key YOUR_API_KEY

If you serve an FP8 checkpoint (e.g. an llm-compressor or ModelOpt export), vLLM auto-detects the weight quantization from config.json; --quantization fp8 is only required to dynamically quantize a BF16 checkpoint at load time. Do not also pass --dtype float8 โ€” that forces the activation dtype, which is a separate setting and is not needed for FP8 weight serving.

Flag explanation:

  • --quantization fp8 โ€” load FP8-quantized weights (W8A8)
  • --kv-cache-dtype fp8 โ€” critical: quantize KV-cache to FP8 instead of defaulting to the model's native FP16
  • --max-model-len 128000 โ€” native context is 131,072 (128K); 128000 leaves a small margin
  • --tensor-parallel-size 1 โ€” single GPU (no TP split)

To check your config's memory needs against the VRAM calculator: Calculator for Llama 3.3 70B @ 128K FP8.

The KV-cache gotcha

vLLM's default kv_cache_dtype is "auto" โ€” which resolves to the model's native precision (FP16 for Llama 3.3). Many deployments set --quantization fp8 for weights and assume KV-cache inherits the savings. It does not โ€” without --kv-cache-dtype fp8, KV-cache stays FP16 and the memory budget reverts to Scenario B.

This is why a configuration that should fit on one H200 can fail at 8โ€“10% GPU memory utilization: the weights are small, but the FP16 KV-cache at 128K context is larger than expected.

Validation steps

  1. Before launch โ€” confirm the checkpoint declares FP8 quantization in its config:
    python -c "import json; qc = json.load(open('config.json')).get('quantization_config'); print(qc)"
    # Expect a dict like {'quant_method': 'fp8', ...}, not None
    
  2. Launch and monitor VRAM:
    vllm serve meta-llama/Llama-3.3-70B-Instruct --quantization fp8 --kv-cache-dtype fp8 ...
    nvidia-smi --query-gpu=memory.used,memory.total --format=csv -l 1
    
  3. Run a 128K-context test โ€” feed 128K tokens and check utilization stays below 110 GB (Scenario A budget):
    curl -X POST http://localhost:8000/v1/chat/completions \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -d '{"model":"meta-llama/Llama-3.3-70B-Instruct","messages":[{"role":"user","content":"<paste 128K tokens of context>"}],"max_tokens":1}'
    
  4. Verify KV-cache dtype in vLLM logs โ€” search for KV cache dtype: fp8.

Concurrency and limitations

  • Batch=1 headroom: at FP8 weights + FP8 KV-cache, 1ร—H200 uses ~101 GB of 141 GB โ€” roughly 40 GB free, enough for about 2 additional 128K-context requests (โ‰ˆ20 GB FP8 KV each) before you approach capacity. Beyond that, scale to 2ร—H200 with tp_size=2.
  • Hopper-only: FP8 W8A8 quantization requires Hopper (H100/H200/B200) or Ada (RTX 4090). On older architectures, fall back to INT4 โ€” see quantization formats guide.
  • Quality trade-off: FP8 W8A8 typically lands within ~1% of BF16 on standard LLM benchmarks (community and vendor reports), but quantized KV-cache can degrade long-context coherence. Validate output quality on your representative prompts โ€” we do not publish a single universal perplexity delta.
  • MLA models: DeepSeek R1 compresses its KV-cache via Multi-Latent Attention (~4 GB at 128K context โ€” small), but its 671B FP8 weights alone are ~671 GB, so single-H200 serving of R1 is not possible at full weight residency; production R1 needs multi-GPU (4ร—H200 INT4 / 6ร—H200 FP8). See the DeepSeek R1 serving guide and FP8 vs INT4 memory analysis.

Common failure modes

SymptomLikely causeFix
OOM at ~110 GB usedKV-cache defaulted to FP16Add --kv-cache-dtype fp8
"Quantization method fp8 not supported"vLLM < 0.6.0 or non-Hopper GPUUpgrade vLLM or use --quantization bitsandbytes with W4A16
Slow first-token at high contextGPU page faults / HBM initializationWarm up with a short prompt before the 128K request
Accuracy degradation >0.5%Aggressive FP8 calibrationUse llm-compressor with dataset calibration instead of no-calibration defaults

Related resources