Which LLMs Can Run on 8GB, 16GB, 24GB, 48GB and 80GB GPUs in 2026?
Which LLMs fit on 8GB, 16GB, 24GB, 48GB and 80GB GPUs — verified fit tables by quantization and context from OpenGPU Radar's VRAM engine.
Direct answer
A model fits a GPU when its weights, KV-cache, runtime overhead, and a 10% fragmentation margin all fit in that GPU's VRAM — at batch 1 and 4,096-token context, per OpenGPU Radar's VRAM calculator using the deterministic VRAM canonical engine (run 2026-10-07). At 8 GB, FP16 reaches only the 1.7B class; INT4 is what gets 8–9B models in the door. At 24 GB everything up to 32B fits at INT4 — but Llama 3.3 70B at INT4 needs 40.8 GB at 4K: it does not fit a single 24 GB GPU. It fits one 48 GB card (40.8, up to 64K); at 80 GB, 70B at FP8 is marginal (79.7).
| VRAM tier | FP16 fits | FP8/INT8 fits | INT4 fits | GPU anchors |
|---|---|---|---|---|
| 8 GB | 2 — up to 1.7B (5.5) | 3 — up to 3.2B (5.3) | 11 — up to 9B-class (6.9) | RTX 4060 8GB (registry) |
| 16 GB | 3 — up to 3.2B (8.8) | 11 — 8B (10.9); 12B marginal (15.2) | 17 — 22B (13.9); 24B marginal (15.1) | T4; RTX 4080 |
| 24 GB | 9 — 8B (20.0); 9B marginal (22.1) | 16 — 14B (17.9) | 21 — 32B (20.1); 70B does not fit (40.8) | RTX 4090; RTX 3090; A10G; RX 7900 XTX |
| 48 GB | 16 — 14B (34.1) | 21 — 32B (38.0) | 26 — 70B (40.8), 72B (42.4) | L40S; RTX 6000 Ada |
| 80 GB | 20 — 32B (72.2); coder-32B marginal (74.2) | 21 — 32B (38.0); 70B marginal (79.7) | 29 — 72B (42.4); mixtral-8x22b marginal (79.5) | A100 80GB; H100 |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU. Totals in GB.
Between tiers? The lower tier is your floor: a 12 GB GPU such as the NVIDIA RTX 3060 12GB fits everything the 8 GB row fits, and the 16 GB row marks the next step up.
What "runs" actually means
"It runs" covers five different situations:
- The weights alone fit. Llama 3.3 70B is 141.2 GB at FP16, 70.6 GB at FP8, 35.3 GB at INT4 — before a token of KV-cache or runtime is loaded.
- The full inference stack fits — weights + KV-cache + 1.2 GB CUDA/runtime + 0.4 GB activation + 10% fragmentation margin, on one GPU. That is the standard below — and what VRAM an LLM actually needs is a memory question, not a parameter-count question.
- It fits only after quantizing — Qwen 2.5 14B: 34.1 GB at FP16 doesn't fit 16 GB; at INT4 its 9.9 GB does.
- It fits with a shorter context — it fits at 4K where it fails at 128K.
- It fits via multi-GPU sharding or CPU/system-RAM offload — real, but the engine models neither; those setups get no numbers here.
Everything below answers claim (2).
How OpenGPU Radar calculates fit
Every number comes from the site's VRAM canonical engine — the calculator's code path:
- Weights = parameters × bytes per parameter (FP16 2.0, FP8 1.0, INT4 0.5 GB per billion parameters).
- KV-cache (GQA) = 2 × layers × KV-heads × head-dim × context × bytes per element. DeepSeek's MLA models use the engine's compressed formula, quantization-independent on the KV side.
- Overheads = + 1.2 GB CUDA/runtime + 0.4 GB activation, then +10% fragmentation headroom on the subtotal.
- Conditions = batch 1, single GPU, inference, default context 4,096 tokens.
- Labels = Fits (headroom ≥ 10% of the requirement), Marginal (< 10%), Does not fit (total > VRAM).
Totals below are weights plus KV plus overhead — the KV-cache grows with context; that's what separates 40.8 GB from 48.1 GB.
Quantization: what each label costs
| Quantization | GB per billion params (weights) | KV bytes per element |
|---|---|---|
| fp16 | 2.0 | 2.0 |
| fp8 | 1.0 | 1.0 |
| int8 | 1.0 | 1.0 |
| int4 | 0.5 | 1.0 |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07. Labels: FP16/BF16 (16 bits), FP8, INT8, INT4/AWQ/GGUF (4 bits).
Three consequences: FP8 and INT8 totals are identical here (one column, two labels); KV does not shrink below 1 byte per element at INT4, so an INT4 model's context cost is half of FP16's, not a quarter; and with no Q5/Q6 modes, GGUF and AWQ map onto 4-bit/8-bit — no Q4_K_M figures exist. Mechanics: the quantization explainer.
8 GB: sub-8B models, INT4 only
FP16 stops at the 1.7B class. INT4 is where 8 GB becomes useful: Llama 3.1 8B needs 6.5 GB, Gemma 2 9B needs 6.9 GB — both fits at 4K.
| Quantization | Fits at 4K (GB) | Marginal |
|---|---|---|
| FP16 | llama-3.2-1b 4.0; smollm2-1.7b 5.5 | — |
| FP8/INT8 | llama-3.2-3b 5.3; llama-3.2-1b 2.9; smollm2-1.7b 3.6 | — |
| INT4 | llama-3.1-8b 6.5; qwen-2.5-coder-7b 6.0; qwen-2.5-7b 6.0; qwen-2.5-vl-7b 6.0; gemma-2-9b 6.9; mistral-7b-v0.3 5.6; mimo-v2-5-free 5.6; glm-4-flash 7.3; llama-3.2-3b 3.5; llama-3.2-1b 2.3; smollm2-1.7b 2.7 | — |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.
Context is the ceiling: Llama 3.1 8B INT4 fits 8 GB at 16K (7.3) but not 32K (8.4). No 8 GB GPU page exists; the registry's "RTX 4060 8GB" (for Llama 3.2 3B) is the anchor.
16 GB: the 7–14B INT4 tier (T4 / RTX 4080 class)
FP16 still stops at 3.2B; 8B runs at FP8 (10.9); INT4 opens the 12–24B range.
Reading the tier tables: each count — like "(17 total)" — is the number of models that Fit at that quantization; "the 11 from 8 GB plus…" means the models that fit the smaller tier also fit here and are counted with them, and the value after each model name is its engine total in GB.
| Quantization | Fits at 4K (GB) | Marginal |
|---|---|---|
| FP16 | llama-3.2-3b 8.8; llama-3.2-1b 4.0; smollm2-1.7b 5.5 | — |
| FP8/INT8 | the 3 from 8 GB plus llama-3.1-8b 10.9; qwen-2.5-7b 10.1; qwen-2.5-coder-7b 10.1; qwen-2.5-vl-7b 10.1; gemma-2-9b 11.9; mistral-7b-v0.3 9.5; mimo-v2-5-free 9.5; glm-4-flash 12.8 (11 total) | mistral-nemo-12b 15.2; gemma-3-12b 15.0 |
| INT4 | the 11 from 8 GB plus qwen-2.5-14b 9.9; qwen-2.5-coder-14b 9.9; codestral-22b 13.9; mistral-nemo-12b 8.6; gemma-3-12b 8.4; phi-4-14b 9.5 (17 total) | mistral-small-24b 15.1 |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.
Codestral 22B fits (13.9); Mistral Small 24B is marginal (15.1). Anchors: NVIDIA T4 (16 GB); registry "RTX 4080 16GB" for Gemma 2 9B.
24 GB: the 32B INT4 tier
FP16 reaches 8B (20.0, 9B marginal at 22.1); FP8 reaches 14B (17.9); INT4 reaches 32B — coder-32b 20.1, r1-distill-qwen-32b 19.4, gemma-2-27b 16.6.
| Quantization | Fits at 4K (GB) | Marginal |
|---|---|---|
| FP16 | the 3 from 16 GB plus llama-3.1-8b 20.0; qwen-2.5-7b 18.5; qwen-2.5-coder-7b 18.5; qwen-2.5-vl-7b 18.5; mistral-7b-v0.3 17.2; mimo-v2-5-free 17.2 (9 total) | gemma-2-9b 22.1; glm-4-flash 23.8 |
| FP8/INT8 | the 11 from 16 GB plus qwen-2.5-14b 17.9; qwen-2.5-coder-14b 17.9; phi-4-14b 17.2; mistral-nemo-12b 15.2; gemma-3-12b 15.0 (16 total) | none at 4K |
| INT4 | the 17 from 16 GB plus deepseek-r1-distill-qwen-32b 19.4; qwen-2.5-coder-32b 20.1; gemma-2-27b 16.6; mistral-small-24b 15.1 (21 total) | none at 4K |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.
70B at INT4 needs 40.8 GB — it does not fit here (answered below). Anchors: RTX 4090, RTX 3090, A10G, RX 7900 XTX.
48 GB: the 70B INT4 tier
FP16 reaches the 14B class (34.1), FP8 reaches 32B (38.0), and INT4 is the tier's reason to exist: the 70/70.6/72.7B class fits from 4K through 64K (40.8–44.4).
| Quantization | Fits at 4K (GB) | Marginal |
|---|---|---|
| FP16 | the 9 from 24 GB plus qwen-2.5-14b 34.1; qwen-2.5-coder-14b 34.1; phi-4-14b 32.6; mistral-nemo-12b 28.6; gemma-3-12b 28.2; glm-4-flash 23.8; gemma-2-9b 22.1 (16 total; 22B+ FP16 does not fit: codestral-22b 50.2) | — |
| FP8/INT8 | the 16 from 24 GB plus qwen-2.5-coder-32b 38.0; deepseek-r1-distill-qwen-32b 37.0; gemma-2-27b 31.5; mistral-small-24b 28.3; codestral-22b 26.0 (21 total; 70B FP8 79.7 does not fit) | — |
| INT4 | the 21 from 24 GB plus llama-3.3-70b 40.8; llama-3.1-70b 41.3; deepseek-r1-distill-llama-70b 40.3; qwen-2.5-vl-72b 41.8; qwen-2.5-72b 42.4 (26 total; 405B INT4 228.8 does not fit) | — |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.
At the full 128K window, 70B INT4 reaches 48.1 GB — just over one card. Anchors: L40S; RTX 6000 Ada (both 48 GB).
80 GB: the datacenter baseline
| Quantization | Fits at 4K (GB) | Marginal |
|---|---|---|
| FP16 | the 16 from 48 GB plus deepseek-r1-distill-qwen-32b 72.2; gemma-2-27b 61.2; mistral-small-24b 54.8; codestral-22b 50.2 (20 total; 70B FP16 157.6 does not fit) | qwen-2.5-coder-32b 74.2 |
| FP8/INT8 | identical set and values to 48 GB (21 total; 70B FP8 = 80.4 at 16K → does not fit) | llama-3.3-70b 79.7; deepseek-r1-distill-llama-70b 78.8 |
| INT4 | the 26 from 48 GB plus command-r-plus 59.1; mistral-large 69.5; ling-3-0-flash 70.1 (29 total; 405B 228.8 does not fit) | mixtral-8x22b 79.5 |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.
70B at FP8: needs 79.7 GB at 4K — Marginal on an 80 GB H100/A100 (0.3 GB headroom), no fit from 16K (80.4 GB), comfortable on 96 GB+. INT4 covers everything through 72B plus the 104B/123B/124B models. Anchors: A100 80GB; H100.
Beyond the tiers (96–288 GB)
- 96 GB: fits 70B FP8 (79.7) and 72B FP8 (82.4); FP16 fits coder-32B (74.2); INT4 fits mixtral-8x22b (79.5).
- 141 GB (H200): fits 70/72B FP8 and everything INT4 up to 141B; 70B FP16 (157.6) does not fit.
- 192 GB (B200 / MI300X): fits 70B FP16 (157.6) and 72B FP16 (163.1); 405B INT4 (228.8) still does not fit.
- 288 GB (B300): fits 405B INT4 (228.8) and command-r-plus FP16 (230.8); mistral-large FP16 (272.6) and ling-3-0-flash FP16 (274.8) are marginal. DeepSeek-V3 INT4 needs 372.3 — multi-GPU only (registry: 8× H200 141GB).
Context length: the hidden multiplier
Quantization sets the floor; context sets the ceiling — same weights, six lengths:
| Model | Quant | 4K | 8K | 16K | 32K | 64K | 128K |
|---|---|---|---|---|---|---|---|
| llama-3.1-8b | FP16 | 20.0 | 20.5 | 21.6 | 23.8 | 28.2 | 36.6 |
| llama-3.1-8b | INT4 | 6.5 | 6.7 | 7.3 | 8.4 | 10.6 | 14.8 |
| llama-3.3-70b | FP16 | 157.6 | 158.0 | 159.0 | 160.9 | 164.8 | 172.1 |
| llama-3.3-70b | INT4 | 40.8 | 41.1 | 41.6 | 42.5 | 44.4 | 48.1 |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.
KV-only rows behind those totals:
| Model | Quant | 4K | 8K | 16K | 32K | 64K | 128K |
|---|---|---|---|---|---|---|---|
| llama-3.1-8b | FP16 | 0.50 | 1.00 | 2.00 | 4.00 | 8.00 | 15.63 |
| llama-3.1-8b | INT4 | 0.25 | 0.50 | 1.00 | 2.00 | 4.00 | 7.81 |
| llama-3.3-70b | FP16 | 0.44 | 0.88 | 1.75 | 3.50 | 7.00 | 13.67 |
| llama-3.3-70b | INT4 | 0.22 | 0.44 | 0.88 | 1.75 | 3.50 | 6.84 |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU. FP8 KV equals INT4 KV — both 1 byte per element.
Two readings: Llama 3.1 8B INT4 grows 6.5 → 14.8 GB from 4K to 128K — the 8 GB card that runs it at 16K (7.3) fails at 32K (8.4). Qwen2.5-Coder 32B FP16 goes 74.2 → 101.2: the gap between one 80 GB GPU (marginal at 4K) and needing 141 GB. Detail: Llama 3.1 8B VRAM.
The 70B-on-24GB question, answered
Llama 3.3 70B (70.6B parameters, 128K window). Weights alone: FP16 141.2 / FP8 70.6 / INT4 35.3 GB.
| Context | FP16 | FP8/INT8 | INT4 |
|---|---|---|---|
| 4,096 | 157.6 | 79.7 | 40.8 |
| 8,192 | 158.0 | 79.9 | 41.1 |
| 16,384 | 159.0 | 80.4 | 41.6 |
| 32,768 | 160.9 | 81.3 | 42.5 |
| 65,536 | 164.8 | 83.3 | 44.4 |
| 128,000 | 172.1 | 86.9 | 48.1 |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.
- Llama 3.3 70B at INT4 needs 40.8 GB at 4K context — a single 24 GB GPU cannot hold it; the repo's GPU evaluator marks 24 GB Insufficient and computes 2 GPUs.
- INT4 fits one 48 GB GPU up to 64K context (44.4 GB); at the full 128K window it reaches 48.1 GB — just over a single 48 GB card.
- FP8 needs 79.7 GB at 4K: Marginal on an 80 GB H100/A100 (0.3 GB headroom), no fit from 16K (80.4 GB), comfortable on 96 GB+.
- FP16 needs 157.6 GB: beyond H200's 141 GB, fits B200/MI300X 192 GB (headroom 34.4 GB).
Fit verdict only — deployment: how to run it locally; sizing: model hub and VRAM breakdown.
FAQ
Does Llama 3.3 70B run on a 24 GB GPU? No. INT4 needs 40.8 GB at 4K; the evaluator computes 2 GPUs. 48 GB covers INT4 up to 64K; 96 GB+ for FP8; 192 GB for FP16.
What is the difference between "fits" and "technically runs"? "Fits" means the full stack — weights + KV-cache + runtime + 10% margin — lives in one GPU's VRAM at batch 1 and 4K context. "Technically runs" usually means sharding or spilling to system RAM. This article only makes the first claim.
Can I offload to CPU or system RAM instead? Valid — but sharding and offload belong to the serving stack (local LLM serving engines); this engine doesn't model them, so no offload figures appear here.
My 8B model fits an 8 GB card at 4K. Why did it stop fitting? Llama 3.1 8B INT4: 6.5 GB at 4K, 7.3 at 16K, 8.4 at 32K — weights never changed; the KV-cache did.
Does any of this apply to Mac or unified memory? No — every figure assumes a single discrete GPU with dedicated VRAM; unified memory isn't modeled.
Limitations and assumptions
- Memory fit only. No throughput, latency, bandwidth, quality, or price claims — pricing: cheapest GPUs.
- Conditions: batch 1, single GPU, inference, default context 4,096; run 2026-10-07. Figures are CALCULATED_ESTIMATES — "needs"/"estimated", never measured. Labels: Fits (headroom ≥ 10% of the requirement), Marginal (< 10%), Does not fit.
- Context windows are the registry's: gemma-2 models 8,000; phi-4 16,000; mistral-7b and codestral 32,000; mixtral-8x22b 65,000; most others 128,000. Runs never exceed 128K.
- KV modeling: exact GQA (high confidence) for llama-3.1-8b, llama-3.3-70b, llama-3.1-70b, llama-3.1-405b, qwen-2.5-coder-32b, qwen-2.5-72b, mistral-nemo-12b, mistral-small-24b, deepseek-v3; modeled estimates (medium confidence) for everything else here — gemma, phi, 7B/14B class, distills, command-r, mistral-large, mixtral, ling. Both are engine output; only the confidence differs.
- Not covered: CPU/RAM offload, multi-GPU throughput, Mac/unified memory, framework versions, quality effects, Q5/Q6 formats.
Related resources
- LLM VRAM Calculator
- GPU requirements chart — same engine, model-first (70B FP8 ≈ 81 GB at 32K matches the 81.3 above)
- What is VRAM? · Quantization explained · KV-cache explained · Local LLM serving engines
- Model hub — sizing pages for every model here
- How to run Llama 3.3 70B locally · Llama 3.3 70B · Llama 3.2 3B
- GPU sizing: T4 · RTX 4090 · L40S · A100 80GB
Related Articles
B200 NVLink 5.0 Scaling: When Does 8x B200 Beat 16x H100?
8x B200 SXM (NVLink 5.0 at 1.8 TB/s) equals 16x H100's aggregate NVLink bandwidth at similar cost — and delivers 50% more 70B inference throughput. Here's the math on bandwidth, power, and cost.
InfrastructureLLM VRAM Calculator: GPU Requirements for Llama, Qwen, DeepSeek & Claude in 2026
Find the right GPU for any LLM: Llama 3.3 70B needs 71 GB of FP8 weights (~81 GB in service), DeepSeek R1 671B needs 786 GB at 128K context. VRAM sizing chart by model.
InfrastructureVerified Free LLM APIs: Which Providers Actually Require No Credit Card in 2026?
Ranked comparison of free LLM APIs requiring no credit card: Google AI Studio (Gemini 2.0 Flash), Groq (Llama 3.3 70B), Cerebras, Cloudflare Workers AI, and more. Verified 2026-09-26.
InfrastructureMinimum Viable Cluster for DeepSeek R1: Sizing Multi-Node Under $15/hr
DeepSeek R1 671B needs 786 GB at FP8/128K (6x H200 minimum, 8x recommended) or 417 GB at INT4 (4x H200). Minimum viable cluster configs and costs under $15/hr.