Infrastructure2026-10-05โ€ขBy Sreeโ€ข5 min read

LLM VRAM Calculator: GPU Requirements for Llama, Qwen, DeepSeek & Claude in 2026

Find the right GPU for any LLM: Llama 3.3 70B needs 71 GB of FP8 weights (~81 GB in service), DeepSeek R1 671B needs 786 GB at 128K context. VRAM sizing chart by model.

Direct answer

GPU sizing starts with weights โ€” parameters ร— bytes per parameter โ€” then adds KV cache, runtime overhead, and headroom. FP8 stores every parameter in 1 byte, so a 70B model is 71 GB of weights (not 140; that is the FP16 figure). At 32K context a 70B FP8 serving instance needs ~81 GB total, which is why it lands on 1ร— H200 or 2ร— H100.

Model sizeFP8 weightsTypical in-service total (32K ctx)Single GPU?Minimum GPU
7Bโ€“8B (Llama 3.1 8B, Qwen 2.5 7B)8 GB~13 GBโœ… RTX 4090, L40S24 GB+
12Bโ€“14B (Gemma 3 12B, Qwen Coder 14B)12โ€“14 GB~17โ€“20 GBโœ… RTX 4090, L40S24 GB+
22Bโ€“27B (Codestral 22B, Gemma 2 27B)22โ€“27 GB~27โ€“33 GBโœ… L40S48 GB
32B (Qwen 2.5 Coder 32B)33 GB~41 GBโœ… L40S, H10048 GB
70B (Llama 3.3 70B, Qwen 2.5 72B)71โ€“73 GB~81โ€“87 GBโŒ (80 GB is marginal)H200 (141 GB) or 2ร— H100
104B (Command R+)104 GB~117 GBโŒ2ร— H100 or 1ร— H200 (141 GB, tight)
141B (Mixtral 8ร—22B)141 GB~158 GBโŒ2ร— H100 (tight) or 2ร— H200
340B (Nemotron-4 340B)340 GB~380 GBโŒ5ร— H100 or 3ร— H200
405B (Llama 3.1 405B)405 GB~482 GB @32KโŒ4ร— H200 (32K) / 5ร— H200 (128K)
671B (DeepSeek R1)671 GB786 GB @128KโŒ6ร— H200 (min) or 5ร— B200

Sources: weights from the deterministic formula (parameters ร— bytes/param) and the model architecture registry (minVramFp8Gb = 71 GB for Llama 70B, 671 GB for DeepSeek R1); totals from the canonical VRAM engine (weights + KV cache + 1.2 GB CUDA overhead + 0.4 GB activations + 10% fragmentation headroom), batch=1, 32K context unless noted. GPU VRAM from the GPU spec database.

How to calculate VRAM requirements

VRAM = weights + KV cache + activations + overhead

Formula:

weights      = params ร— bytes_per_param        (FP8 = 1, FP16 = 2, INT4 = 0.5)
kv_cache     = 2 ร— layers ร— kv_heads ร— head_dim ร— context ร— bytes_per_kv_element
total        = (weights + kv_cache + 1.2 GB CUDA + 0.4 GB activations) ร— 1.10

The 10% headroom covers allocator fragmentation โ€” serving frameworks routinely reserve it, and skipping it is how "fits on paper" deployments OOM.

Quantization formats

FormatBytes/paramWeights vs FP16Quality impactRecommended for
FP16/BF162baselinenonetraining, accuracy-critical eval
FP81โˆ’50%negligible in practice for servingproduction serving
INT81โˆ’50%negligiblesupported on older hardware
INT4 (AWQ/GGUF)0.5โˆ’75%measurable โ€” benchmark on your taskmemory-constrained deployment

Note: perplexity deltas from quantization are model- and dataset-dependent; treat any single published percentage as an estimate, not a guarantee. See the DeepSeek R1 FP8 vs INT4 analysis for the measurement approach.

KV cache calculation

KV cache (GB) = 2 ร— layers ร— kv_heads ร— head_dim ร— context_length ร— bytes_per_element รท 1,073,741,824

Worked examples at 128K context, FP8 KV (1 byte/element), batch=1, using the architecture presets in our VRAM engine:

ModelLayers (engine preset)KV @128K
Llama 3.1 8B328 GB
Qwen 2.5 Coder 32B5213 GB
Llama 3.3 70B28 (see note)7 GB โ€” treat as 7โ€“20 GB
Qwen 2.5 72B8020 GB
Llama 3.1 405B126126 GB
DeepSeek R1 (MLA)compressed KV42 GB

Note on Llama 70B layers: our registry lists 28 layers for Llama 3.3 70B and 80 for Llama 3.1 70B โ€” the same architecture. Until that record is reconciled, size 70B KV cache against the 80-layer figure (20 GB @128K) rather than the low one. DeepSeek R1 uses MLA compressed KV (~0.5 KB/token/layer), which is why its KV cache is only 42 GB despite 671B parameters.

VRAM chart by model

ModelParamsFP8 weightsINT4 weightsTotal FP8 @32KNeeds GPU
Llama 3.1 8B8.03B8 GB4 GB~13 GBRTX 4090 (24 GB) โœ…
Qwen 2.5 7B7.6B8 GB4 GB~13 GBRTX 4090 (24 GB) โœ…
Gemma 3 12B12B12 GB6 GB~17 GBRTX 4090, L40S โœ…
Codestral 22B22B22 GB11 GB~27 GBL40S (48 GB) โœ…
Qwen 2.5 Coder 32B32.5B33 GB16 GB~41 GBL40S, H100 โœ…
Llama 3.3 70B70.6B71 GB35 GB~81 GBH200 (141 GB) โœ…; 2ร— H100 โœ…
Qwen 2.5 72B72.7B73 GB36 GB~87 GBH200 โœ…
Mistral Small 24B24B24 GB12 GB~30 GBL40S โœ…
Mistral Large 123B123B123 GB62 GB~138 GB2ร— H100 โœ…; 1ร— H200 (tight) โœ…
Command R+ 104B104B104 GB52 GB~117 GB2ร— H100 โœ…; 1ร— H200 โœ…
Mixtral 8ร—22B141B141 GB70 GB~158 GB2ร— H100 โœ… (tight)
Llama 3.1 405B405B405 GB203 GB~482 GB @32K4ร— H200 โœ… (32K); 5ร— H200 @128K
DeepSeek R1 671B671B671 GB336 GB786 GB @128K6ร— H200 โœ… (min); 8ร— H200 comfortable

Key insight: INT4 puts 70B weights in 35 GB (~43 GB in service at 32K) โ€” comfortable on a single H100. FP8 at 71 GB weights / ~81 GB in service is what pushes 70B onto H200.

GPU VRAM reference

GPUVRAMFP8 TFLOPSCost (Vast.ai, 2026-10-03)Recommended use case
RTX 409024 GB165$0.34/hr7Bโ€“8B FP8/INT4, prototyping
L40S48 GB733$0.69/hrup to ~32B FP8, batch inference
H100 SXM580 GB1979$1.89/hr70B INT4 single-GPU; 32B FP8
H200 SXM5141 GB1979$2.79/hr70Bโ€“120B FP8 single-GPU
B200 SXM192 GB2250$3.99/hr123Bโ€“190B FP8; 671B INT4 (5ร—)

Sources: GPU specs from the GPU spec database; pricing from providers.json (observed 2026-10-03).

Model-by-model VRAM guide

Llama 3.3 70B (70.6B parameters)

  • FP8: 71 GB weights; ~81 GB total at 32K context
  • FP16: 141 GB weights; ~161 GB total at 32K
  • INT4: 35 GB weights; ~43 GB total at 32K
  • Recommended GPU: 1ร— H200 (141 GB) โœ… or 2ร— H100 โœ… โ€” FP8 does not fit a single 80 GB H100 once KV cache and overhead are included
  • Free API: Groq (Llama 3.3 70B, 30 RPM, no credit card)
  • Calculator: Llama 3.3 70B VRAM

DeepSeek R1 671B (671B parameters, 37B active MoE)

  • FP8: 671 GB weights; 786 GB total at 128K context (MLA KV cache = 42 GB)
  • INT4: 336 GB weights; 417 GB total at 128K
  • FP16: 1,342 GB weights; 1,524 GB total at 128K
  • Recommended GPU: 8ร— H200 (1,128 GB) โœ… for FP8; 5ร— B200 (960 GB) โœ… for FP8; 3ร— H200 (423 GB) โœ… or 4ร— B200 (768 GB) โœ… for INT4
  • API pricing: DeepSeek API ($0.55/M in, $2.19/M out)
  • Calculator: DeepSeek R1 VRAM

Llama 3.1 8B (8.03B parameters)

  • FP8: 8 GB weights; ~13 GB total at 32K, ~19 GB at 128K
  • INT4: 4 GB weights; ~8 GB total at 32K
  • Recommended GPU: RTX 4090 (24 GB) โœ… โ€” fits on a single consumer GPU
  • Free API: Groq (30 RPM, 14,400 requests/day)
  • Calculator: Llama 3.1 8B VRAM

Qwen 2.5 Coder 32B (32.5B parameters)

  • FP8: 33 GB weights; ~41 GB total at 32K, ~52 GB at 128K
  • INT4: 16 GB weights; ~23 GB total at 32K
  • Recommended GPU: 1ร— L40S (48 GB) โœ… or 1ร— H100 โœ…
  • Free API: Groq (Qwen 2.5 Coder 32B, 30 RPM โ€” listed in our free API dataset)
  • Calculator: Qwen 2.5 Coder 32B VRAM

Claude 3.5 Sonnet (Anthropic, API only)

  • Weights: closed โ€” no VRAM sizing is possible; there are no verifiable parameter counts to publish
  • API pricing: $3/M input, $15/M output
  • Free tier: not available โ€” requires credit card

128K context VRAM

For long-context workloads the KV cache stops being negligible. FP8 totals at 128K context, batch=1:

ModelFP8 weightsKV cache (128K)Total FP8 @128K
7Bโ€“8B8 GB8 GB~19 GB
22Bโ€“32B22โ€“33 GB~7โ€“13 GB~30โ€“52 GB
70B71 GB7โ€“20 GB (layer record note above)~87โ€“101 GB
405B405 GB126 GB~586 GB
671B (MoE, MLA)671 GB42 GB786 GB

Methodology: canonical VRAM engine, FP8 KV elements (1 byte), architecture presets as documented per row. The old version of this table listed FP16 weight figures under an FP8 header and understated KV cache by an order of magnitude โ€” both corrected.

What changes the result?

  1. Quantization: INT4 cuts 70B weights to 35 GB โ€” a single H100 goes from โŒ to โœ…
  2. PagedAttention / KV offloading: lets you page KV cache to CPU memory โ€” functional on 80 GB GPUs, at a throughput cost
  3. Tensor parallelism: 2ร— H100 = 160 GB aggregate for 70B; 8ร— H200 = 1,128 GB for 671B
  4. Batching: larger batches multiply KV cache (per-sequence) and activation memory โ€” size headroom per target batch, not just per single stream
  5. MoE models: 671B MoE activates only 37B parameters per token but must resident-load all experts โ€” 671 GB at FP8 regardless of active count

Alternatives

Conclusion

In-service VRAM (FP8, 32K)Suitable modelsMinimum GPUCost/hr (Vast.ai)
โ‰ค 24 GB7Bโ€“14BRTX 4090$0.34
โ‰ค 48 GBup to ~32BL40S$0.69
โ‰ค 80 GB70B in INT4; โ‰ค32B FP8H100$1.89
โ‰ค 141 GB70Bโ€“120B FP8H200$2.79
โ‰ค 192 GBup to ~190B FP8B200$3.99
> 192 GB405Bโ€“671B+multi-GPU (8ร— H200 = $22.32)โ€”

Calculate exact VRAM for your model and context window with the GPU VRAM Calculator, or browse model pages for per-model GPU compatibility data.

Related resources