LLM VRAM Calculator: GPU Requirements for Llama, Qwen, DeepSeek & Claude in 2026
Find the right GPU for any LLM: Llama 3.3 70B needs 71 GB of FP8 weights (~81 GB in service), DeepSeek R1 671B needs 786 GB at 128K context. VRAM sizing chart by model.
Direct answer
GPU sizing starts with weights โ parameters ร bytes per parameter โ then adds KV cache, runtime overhead, and headroom. FP8 stores every parameter in 1 byte, so a 70B model is 71 GB of weights (not 140; that is the FP16 figure). At 32K context a 70B FP8 serving instance needs ~81 GB total, which is why it lands on 1ร H200 or 2ร H100.
| Model size | FP8 weights | Typical in-service total (32K ctx) | Single GPU? | Minimum GPU |
|---|---|---|---|---|
| 7Bโ8B (Llama 3.1 8B, Qwen 2.5 7B) | 8 GB | ~13 GB | โ RTX 4090, L40S | 24 GB+ |
| 12Bโ14B (Gemma 3 12B, Qwen Coder 14B) | 12โ14 GB | ~17โ20 GB | โ RTX 4090, L40S | 24 GB+ |
| 22Bโ27B (Codestral 22B, Gemma 2 27B) | 22โ27 GB | ~27โ33 GB | โ L40S | 48 GB |
| 32B (Qwen 2.5 Coder 32B) | 33 GB | ~41 GB | โ L40S, H100 | 48 GB |
| 70B (Llama 3.3 70B, Qwen 2.5 72B) | 71โ73 GB | ~81โ87 GB | โ (80 GB is marginal) | H200 (141 GB) or 2ร H100 |
| 104B (Command R+) | 104 GB | ~117 GB | โ | 2ร H100 or 1ร H200 (141 GB, tight) |
| 141B (Mixtral 8ร22B) | 141 GB | ~158 GB | โ | 2ร H100 (tight) or 2ร H200 |
| 340B (Nemotron-4 340B) | 340 GB | ~380 GB | โ | 5ร H100 or 3ร H200 |
| 405B (Llama 3.1 405B) | 405 GB | ~482 GB @32K | โ | 4ร H200 (32K) / 5ร H200 (128K) |
| 671B (DeepSeek R1) | 671 GB | 786 GB @128K | โ | 6ร H200 (min) or 5ร B200 |
Sources: weights from the deterministic formula (parameters ร bytes/param) and the model architecture registry (
minVramFp8Gb= 71 GB for Llama 70B, 671 GB for DeepSeek R1); totals from the canonical VRAM engine (weights + KV cache + 1.2 GB CUDA overhead + 0.4 GB activations + 10% fragmentation headroom), batch=1, 32K context unless noted. GPU VRAM from the GPU spec database.
How to calculate VRAM requirements
VRAM = weights + KV cache + activations + overhead
Formula:
weights = params ร bytes_per_param (FP8 = 1, FP16 = 2, INT4 = 0.5)
kv_cache = 2 ร layers ร kv_heads ร head_dim ร context ร bytes_per_kv_element
total = (weights + kv_cache + 1.2 GB CUDA + 0.4 GB activations) ร 1.10
The 10% headroom covers allocator fragmentation โ serving frameworks routinely reserve it, and skipping it is how "fits on paper" deployments OOM.
Quantization formats
| Format | Bytes/param | Weights vs FP16 | Quality impact | Recommended for |
|---|---|---|---|---|
| FP16/BF16 | 2 | baseline | none | training, accuracy-critical eval |
| FP8 | 1 | โ50% | negligible in practice for serving | production serving |
| INT8 | 1 | โ50% | negligible | supported on older hardware |
| INT4 (AWQ/GGUF) | 0.5 | โ75% | measurable โ benchmark on your task | memory-constrained deployment |
Note: perplexity deltas from quantization are model- and dataset-dependent; treat any single published percentage as an estimate, not a guarantee. See the DeepSeek R1 FP8 vs INT4 analysis for the measurement approach.
KV cache calculation
KV cache (GB) = 2 ร layers ร kv_heads ร head_dim ร context_length ร bytes_per_element รท 1,073,741,824
Worked examples at 128K context, FP8 KV (1 byte/element), batch=1, using the architecture presets in our VRAM engine:
| Model | Layers (engine preset) | KV @128K |
|---|---|---|
| Llama 3.1 8B | 32 | 8 GB |
| Qwen 2.5 Coder 32B | 52 | 13 GB |
| Llama 3.3 70B | 28 (see note) | 7 GB โ treat as 7โ20 GB |
| Qwen 2.5 72B | 80 | 20 GB |
| Llama 3.1 405B | 126 | 126 GB |
| DeepSeek R1 (MLA) | compressed KV | 42 GB |
Note on Llama 70B layers: our registry lists 28 layers for Llama 3.3 70B and 80 for Llama 3.1 70B โ the same architecture. Until that record is reconciled, size 70B KV cache against the 80-layer figure (20 GB @128K) rather than the low one. DeepSeek R1 uses MLA compressed KV (~0.5 KB/token/layer), which is why its KV cache is only 42 GB despite 671B parameters.
VRAM chart by model
| Model | Params | FP8 weights | INT4 weights | Total FP8 @32K | Needs GPU |
|---|---|---|---|---|---|
| Llama 3.1 8B | 8.03B | 8 GB | 4 GB | ~13 GB | RTX 4090 (24 GB) โ |
| Qwen 2.5 7B | 7.6B | 8 GB | 4 GB | ~13 GB | RTX 4090 (24 GB) โ |
| Gemma 3 12B | 12B | 12 GB | 6 GB | ~17 GB | RTX 4090, L40S โ |
| Codestral 22B | 22B | 22 GB | 11 GB | ~27 GB | L40S (48 GB) โ |
| Qwen 2.5 Coder 32B | 32.5B | 33 GB | 16 GB | ~41 GB | L40S, H100 โ |
| Llama 3.3 70B | 70.6B | 71 GB | 35 GB | ~81 GB | H200 (141 GB) โ ; 2ร H100 โ |
| Qwen 2.5 72B | 72.7B | 73 GB | 36 GB | ~87 GB | H200 โ |
| Mistral Small 24B | 24B | 24 GB | 12 GB | ~30 GB | L40S โ |
| Mistral Large 123B | 123B | 123 GB | 62 GB | ~138 GB | 2ร H100 โ ; 1ร H200 (tight) โ |
| Command R+ 104B | 104B | 104 GB | 52 GB | ~117 GB | 2ร H100 โ ; 1ร H200 โ |
| Mixtral 8ร22B | 141B | 141 GB | 70 GB | ~158 GB | 2ร H100 โ (tight) |
| Llama 3.1 405B | 405B | 405 GB | 203 GB | ~482 GB @32K | 4ร H200 โ (32K); 5ร H200 @128K |
| DeepSeek R1 671B | 671B | 671 GB | 336 GB | 786 GB @128K | 6ร H200 โ (min); 8ร H200 comfortable |
Key insight: INT4 puts 70B weights in 35 GB (~43 GB in service at 32K) โ comfortable on a single H100. FP8 at 71 GB weights / ~81 GB in service is what pushes 70B onto H200.
GPU VRAM reference
| GPU | VRAM | FP8 TFLOPS | Cost (Vast.ai, 2026-10-03) | Recommended use case |
|---|---|---|---|---|
| RTX 4090 | 24 GB | 165 | $0.34/hr | 7Bโ8B FP8/INT4, prototyping |
| L40S | 48 GB | 733 | $0.69/hr | up to ~32B FP8, batch inference |
| H100 SXM5 | 80 GB | 1979 | $1.89/hr | 70B INT4 single-GPU; 32B FP8 |
| H200 SXM5 | 141 GB | 1979 | $2.79/hr | 70Bโ120B FP8 single-GPU |
| B200 SXM | 192 GB | 2250 | $3.99/hr | 123Bโ190B FP8; 671B INT4 (5ร) |
Sources: GPU specs from the GPU spec database; pricing from
providers.json(observed 2026-10-03).
Model-by-model VRAM guide
Llama 3.3 70B (70.6B parameters)
- FP8: 71 GB weights; ~81 GB total at 32K context
- FP16: 141 GB weights; ~161 GB total at 32K
- INT4: 35 GB weights; ~43 GB total at 32K
- Recommended GPU: 1ร H200 (141 GB) โ or 2ร H100 โ โ FP8 does not fit a single 80 GB H100 once KV cache and overhead are included
- Free API: Groq (Llama 3.3 70B, 30 RPM, no credit card)
- Calculator: Llama 3.3 70B VRAM
DeepSeek R1 671B (671B parameters, 37B active MoE)
- FP8: 671 GB weights; 786 GB total at 128K context (MLA KV cache = 42 GB)
- INT4: 336 GB weights; 417 GB total at 128K
- FP16: 1,342 GB weights; 1,524 GB total at 128K
- Recommended GPU: 8ร H200 (1,128 GB) โ for FP8; 5ร B200 (960 GB) โ for FP8; 3ร H200 (423 GB) โ or 4ร B200 (768 GB) โ for INT4
- API pricing: DeepSeek API ($0.55/M in, $2.19/M out)
- Calculator: DeepSeek R1 VRAM
Llama 3.1 8B (8.03B parameters)
- FP8: 8 GB weights; ~13 GB total at 32K, ~19 GB at 128K
- INT4: 4 GB weights; ~8 GB total at 32K
- Recommended GPU: RTX 4090 (24 GB) โ โ fits on a single consumer GPU
- Free API: Groq (30 RPM, 14,400 requests/day)
- Calculator: Llama 3.1 8B VRAM
Qwen 2.5 Coder 32B (32.5B parameters)
- FP8: 33 GB weights; ~41 GB total at 32K, ~52 GB at 128K
- INT4: 16 GB weights; ~23 GB total at 32K
- Recommended GPU: 1ร L40S (48 GB) โ or 1ร H100 โ
- Free API: Groq (Qwen 2.5 Coder 32B, 30 RPM โ listed in our free API dataset)
- Calculator: Qwen 2.5 Coder 32B VRAM
Claude 3.5 Sonnet (Anthropic, API only)
- Weights: closed โ no VRAM sizing is possible; there are no verifiable parameter counts to publish
- API pricing: $3/M input, $15/M output
- Free tier: not available โ requires credit card
128K context VRAM
For long-context workloads the KV cache stops being negligible. FP8 totals at 128K context, batch=1:
| Model | FP8 weights | KV cache (128K) | Total FP8 @128K |
|---|---|---|---|
| 7Bโ8B | 8 GB | 8 GB | ~19 GB |
| 22Bโ32B | 22โ33 GB | ~7โ13 GB | ~30โ52 GB |
| 70B | 71 GB | 7โ20 GB (layer record note above) | ~87โ101 GB |
| 405B | 405 GB | 126 GB | ~586 GB |
| 671B (MoE, MLA) | 671 GB | 42 GB | 786 GB |
Methodology: canonical VRAM engine, FP8 KV elements (1 byte), architecture presets as documented per row. The old version of this table listed FP16 weight figures under an FP8 header and understated KV cache by an order of magnitude โ both corrected.
What changes the result?
- Quantization: INT4 cuts 70B weights to 35 GB โ a single H100 goes from โ to โ
- PagedAttention / KV offloading: lets you page KV cache to CPU memory โ functional on 80 GB GPUs, at a throughput cost
- Tensor parallelism: 2ร H100 = 160 GB aggregate for 70B; 8ร H200 = 1,128 GB for 671B
- Batching: larger batches multiply KV cache (per-sequence) and activation memory โ size headroom per target batch, not just per single stream
- MoE models: 671B MoE activates only 37B parameters per token but must resident-load all experts โ 671 GB at FP8 regardless of active count
Alternatives
- For 70B models: H200 ($2.79/hr) vs B200 ($3.99/hr) โ compare GPUs
- For 32B models: L40S ($0.69/hr) runs 32B FP8 in ~41 GB โ check L40S pricing
- For prototyping: RTX 4090 ($0.34/hr) for 7Bโ14B โ check RTX 4090 pricing
- For production 671B+: B200 cluster โ B200 pricing
Conclusion
| In-service VRAM (FP8, 32K) | Suitable models | Minimum GPU | Cost/hr (Vast.ai) |
|---|---|---|---|
| โค 24 GB | 7Bโ14B | RTX 4090 | $0.34 |
| โค 48 GB | up to ~32B | L40S | $0.69 |
| โค 80 GB | 70B in INT4; โค32B FP8 | H100 | $1.89 |
| โค 141 GB | 70Bโ120B FP8 | H200 | $2.79 |
| โค 192 GB | up to ~190B FP8 | B200 | $3.99 |
| > 192 GB | 405Bโ671B+ | multi-GPU (8ร H200 = $22.32) | โ |
Calculate exact VRAM for your model and context window with the GPU VRAM Calculator, or browse model pages for per-model GPU compatibility data.
Related resources
- Inference Cost Calculator โ tokens/sec per $/hr by GPU and model
- Llama 3.3 70B Model Page โ VRAM, API pricing, GPU match
- DeepSeek R1 Model Page โ 671B MoE VRAM requirements
- H200 Cloud Pricing โ live rates for the 141 GB GPU
- B200 Blackwell Pricing โ rates for the 192 GB GPU
- Free LLM APIs (No Credit Card) โ alternative to self-hosting
- What is VRAM? โ educational guide on GPU memory and LLM requirements
Related Articles
B200 NVLink 5.0 Scaling: When Does 8x B200 Beat 16x H100?
8x B200 SXM (NVLink 5.0 at 1.8 TB/s) equals 16x H100's aggregate NVLink bandwidth at similar cost โ and delivers 50% more 70B inference throughput. Here's the math on bandwidth, power, and cost.
InfrastructureVerified Free LLM APIs: Which Providers Actually Require No Credit Card in 2026?
Ranked comparison of free LLM APIs requiring no credit card: Google AI Studio (Gemini 2.0 Flash), Groq (Llama 3.3 70B), Cerebras, Cloudflare Workers AI, and more. Verified 2026-09-26.
InfrastructureMinimum Viable Cluster for DeepSeek R1: Sizing Multi-Node Under $15/hr
DeepSeek R1 671B needs 786 GB at FP8/128K (6x H200 minimum, 8x recommended) or 417 GB at INT4 (4x H200). Minimum viable cluster configs and costs under $15/hr.