Infrastructure2026-10-07•By Sree•5 min read

Which LLMs Can Run on 8GB, 16GB, 24GB, 48GB and 80GB GPUs in 2026?

Which LLMs fit on 8GB, 16GB, 24GB, 48GB and 80GB GPUs — verified fit tables by quantization and context from OpenGPU Radar's VRAM engine.

Direct answer

A model fits a GPU when its weights, KV-cache, runtime overhead, and a 10% fragmentation margin all fit in that GPU's VRAM — at batch 1 and 4,096-token context, per OpenGPU Radar's VRAM calculator using the deterministic VRAM canonical engine (run 2026-10-07). At 8 GB, FP16 reaches only the 1.7B class; INT4 is what gets 8–9B models in the door. At 24 GB everything up to 32B fits at INT4 — but Llama 3.3 70B at INT4 needs 40.8 GB at 4K: it does not fit a single 24 GB GPU. It fits one 48 GB card (40.8, up to 64K); at 80 GB, 70B at FP8 is marginal (79.7).

VRAM tierFP16 fitsFP8/INT8 fitsINT4 fitsGPU anchors
8 GB2 — up to 1.7B (5.5)3 — up to 3.2B (5.3)11 — up to 9B-class (6.9)RTX 4060 8GB (registry)
16 GB3 — up to 3.2B (8.8)11 — 8B (10.9); 12B marginal (15.2)17 — 22B (13.9); 24B marginal (15.1)T4; RTX 4080
24 GB9 — 8B (20.0); 9B marginal (22.1)16 — 14B (17.9)21 — 32B (20.1); 70B does not fit (40.8)RTX 4090; RTX 3090; A10G; RX 7900 XTX
48 GB16 — 14B (34.1)21 — 32B (38.0)26 — 70B (40.8), 72B (42.4)L40S; RTX 6000 Ada
80 GB20 — 32B (72.2); coder-32B marginal (74.2)21 — 32B (38.0); 70B marginal (79.7)29 — 72B (42.4); mixtral-8x22b marginal (79.5)A100 80GB; H100

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU. Totals in GB.

Between tiers? The lower tier is your floor: a 12 GB GPU such as the NVIDIA RTX 3060 12GB fits everything the 8 GB row fits, and the 16 GB row marks the next step up.

What "runs" actually means

"It runs" covers five different situations:

  1. The weights alone fit. Llama 3.3 70B is 141.2 GB at FP16, 70.6 GB at FP8, 35.3 GB at INT4 — before a token of KV-cache or runtime is loaded.
  2. The full inference stack fits — weights + KV-cache + 1.2 GB CUDA/runtime + 0.4 GB activation + 10% fragmentation margin, on one GPU. That is the standard below — and what VRAM an LLM actually needs is a memory question, not a parameter-count question.
  3. It fits only after quantizing — Qwen 2.5 14B: 34.1 GB at FP16 doesn't fit 16 GB; at INT4 its 9.9 GB does.
  4. It fits with a shorter context — it fits at 4K where it fails at 128K.
  5. It fits via multi-GPU sharding or CPU/system-RAM offload — real, but the engine models neither; those setups get no numbers here.

Everything below answers claim (2).

How OpenGPU Radar calculates fit

Every number comes from the site's VRAM canonical engine — the calculator's code path:

  • Weights = parameters × bytes per parameter (FP16 2.0, FP8 1.0, INT4 0.5 GB per billion parameters).
  • KV-cache (GQA) = 2 × layers × KV-heads × head-dim × context × bytes per element. DeepSeek's MLA models use the engine's compressed formula, quantization-independent on the KV side.
  • Overheads = + 1.2 GB CUDA/runtime + 0.4 GB activation, then +10% fragmentation headroom on the subtotal.
  • Conditions = batch 1, single GPU, inference, default context 4,096 tokens.
  • Labels = Fits (headroom ≥ 10% of the requirement), Marginal (< 10%), Does not fit (total > VRAM).

Totals below are weights plus KV plus overhead — the KV-cache grows with context; that's what separates 40.8 GB from 48.1 GB.

Quantization: what each label costs

QuantizationGB per billion params (weights)KV bytes per element
fp162.02.0
fp81.01.0
int81.01.0
int40.51.0

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07. Labels: FP16/BF16 (16 bits), FP8, INT8, INT4/AWQ/GGUF (4 bits).

Three consequences: FP8 and INT8 totals are identical here (one column, two labels); KV does not shrink below 1 byte per element at INT4, so an INT4 model's context cost is half of FP16's, not a quarter; and with no Q5/Q6 modes, GGUF and AWQ map onto 4-bit/8-bit — no Q4_K_M figures exist. Mechanics: the quantization explainer.

8 GB: sub-8B models, INT4 only

FP16 stops at the 1.7B class. INT4 is where 8 GB becomes useful: Llama 3.1 8B needs 6.5 GB, Gemma 2 9B needs 6.9 GB — both fits at 4K.

QuantizationFits at 4K (GB)Marginal
FP16llama-3.2-1b 4.0; smollm2-1.7b 5.5—
FP8/INT8llama-3.2-3b 5.3; llama-3.2-1b 2.9; smollm2-1.7b 3.6—
INT4llama-3.1-8b 6.5; qwen-2.5-coder-7b 6.0; qwen-2.5-7b 6.0; qwen-2.5-vl-7b 6.0; gemma-2-9b 6.9; mistral-7b-v0.3 5.6; mimo-v2-5-free 5.6; glm-4-flash 7.3; llama-3.2-3b 3.5; llama-3.2-1b 2.3; smollm2-1.7b 2.7—

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.

Context is the ceiling: Llama 3.1 8B INT4 fits 8 GB at 16K (7.3) but not 32K (8.4). No 8 GB GPU page exists; the registry's "RTX 4060 8GB" (for Llama 3.2 3B) is the anchor.

16 GB: the 7–14B INT4 tier (T4 / RTX 4080 class)

FP16 still stops at 3.2B; 8B runs at FP8 (10.9); INT4 opens the 12–24B range.

Reading the tier tables: each count — like "(17 total)" — is the number of models that Fit at that quantization; "the 11 from 8 GB plus…" means the models that fit the smaller tier also fit here and are counted with them, and the value after each model name is its engine total in GB.

QuantizationFits at 4K (GB)Marginal
FP16llama-3.2-3b 8.8; llama-3.2-1b 4.0; smollm2-1.7b 5.5—
FP8/INT8the 3 from 8 GB plus llama-3.1-8b 10.9; qwen-2.5-7b 10.1; qwen-2.5-coder-7b 10.1; qwen-2.5-vl-7b 10.1; gemma-2-9b 11.9; mistral-7b-v0.3 9.5; mimo-v2-5-free 9.5; glm-4-flash 12.8 (11 total)mistral-nemo-12b 15.2; gemma-3-12b 15.0
INT4the 11 from 8 GB plus qwen-2.5-14b 9.9; qwen-2.5-coder-14b 9.9; codestral-22b 13.9; mistral-nemo-12b 8.6; gemma-3-12b 8.4; phi-4-14b 9.5 (17 total)mistral-small-24b 15.1

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.

Codestral 22B fits (13.9); Mistral Small 24B is marginal (15.1). Anchors: NVIDIA T4 (16 GB); registry "RTX 4080 16GB" for Gemma 2 9B.

24 GB: the 32B INT4 tier

FP16 reaches 8B (20.0, 9B marginal at 22.1); FP8 reaches 14B (17.9); INT4 reaches 32B — coder-32b 20.1, r1-distill-qwen-32b 19.4, gemma-2-27b 16.6.

QuantizationFits at 4K (GB)Marginal
FP16the 3 from 16 GB plus llama-3.1-8b 20.0; qwen-2.5-7b 18.5; qwen-2.5-coder-7b 18.5; qwen-2.5-vl-7b 18.5; mistral-7b-v0.3 17.2; mimo-v2-5-free 17.2 (9 total)gemma-2-9b 22.1; glm-4-flash 23.8
FP8/INT8the 11 from 16 GB plus qwen-2.5-14b 17.9; qwen-2.5-coder-14b 17.9; phi-4-14b 17.2; mistral-nemo-12b 15.2; gemma-3-12b 15.0 (16 total)none at 4K
INT4the 17 from 16 GB plus deepseek-r1-distill-qwen-32b 19.4; qwen-2.5-coder-32b 20.1; gemma-2-27b 16.6; mistral-small-24b 15.1 (21 total)none at 4K

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.

70B at INT4 needs 40.8 GB — it does not fit here (answered below). Anchors: RTX 4090, RTX 3090, A10G, RX 7900 XTX.

48 GB: the 70B INT4 tier

FP16 reaches the 14B class (34.1), FP8 reaches 32B (38.0), and INT4 is the tier's reason to exist: the 70/70.6/72.7B class fits from 4K through 64K (40.8–44.4).

QuantizationFits at 4K (GB)Marginal
FP16the 9 from 24 GB plus qwen-2.5-14b 34.1; qwen-2.5-coder-14b 34.1; phi-4-14b 32.6; mistral-nemo-12b 28.6; gemma-3-12b 28.2; glm-4-flash 23.8; gemma-2-9b 22.1 (16 total; 22B+ FP16 does not fit: codestral-22b 50.2)—
FP8/INT8the 16 from 24 GB plus qwen-2.5-coder-32b 38.0; deepseek-r1-distill-qwen-32b 37.0; gemma-2-27b 31.5; mistral-small-24b 28.3; codestral-22b 26.0 (21 total; 70B FP8 79.7 does not fit)—
INT4the 21 from 24 GB plus llama-3.3-70b 40.8; llama-3.1-70b 41.3; deepseek-r1-distill-llama-70b 40.3; qwen-2.5-vl-72b 41.8; qwen-2.5-72b 42.4 (26 total; 405B INT4 228.8 does not fit)—

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.

At the full 128K window, 70B INT4 reaches 48.1 GB — just over one card. Anchors: L40S; RTX 6000 Ada (both 48 GB).

80 GB: the datacenter baseline

QuantizationFits at 4K (GB)Marginal
FP16the 16 from 48 GB plus deepseek-r1-distill-qwen-32b 72.2; gemma-2-27b 61.2; mistral-small-24b 54.8; codestral-22b 50.2 (20 total; 70B FP16 157.6 does not fit)qwen-2.5-coder-32b 74.2
FP8/INT8identical set and values to 48 GB (21 total; 70B FP8 = 80.4 at 16K → does not fit)llama-3.3-70b 79.7; deepseek-r1-distill-llama-70b 78.8
INT4the 26 from 48 GB plus command-r-plus 59.1; mistral-large 69.5; ling-3-0-flash 70.1 (29 total; 405B 228.8 does not fit)mixtral-8x22b 79.5

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.

70B at FP8: needs 79.7 GB at 4K — Marginal on an 80 GB H100/A100 (0.3 GB headroom), no fit from 16K (80.4 GB), comfortable on 96 GB+. INT4 covers everything through 72B plus the 104B/123B/124B models. Anchors: A100 80GB; H100.

Beyond the tiers (96–288 GB)

  • 96 GB: fits 70B FP8 (79.7) and 72B FP8 (82.4); FP16 fits coder-32B (74.2); INT4 fits mixtral-8x22b (79.5).
  • 141 GB (H200): fits 70/72B FP8 and everything INT4 up to 141B; 70B FP16 (157.6) does not fit.
  • 192 GB (B200 / MI300X): fits 70B FP16 (157.6) and 72B FP16 (163.1); 405B INT4 (228.8) still does not fit.
  • 288 GB (B300): fits 405B INT4 (228.8) and command-r-plus FP16 (230.8); mistral-large FP16 (272.6) and ling-3-0-flash FP16 (274.8) are marginal. DeepSeek-V3 INT4 needs 372.3 — multi-GPU only (registry: 8× H200 141GB).

Context length: the hidden multiplier

Quantization sets the floor; context sets the ceiling — same weights, six lengths:

ModelQuant4K8K16K32K64K128K
llama-3.1-8bFP1620.020.521.623.828.236.6
llama-3.1-8bINT46.56.77.38.410.614.8
llama-3.3-70bFP16157.6158.0159.0160.9164.8172.1
llama-3.3-70bINT440.841.141.642.544.448.1

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.

KV-only rows behind those totals:

ModelQuant4K8K16K32K64K128K
llama-3.1-8bFP160.501.002.004.008.0015.63
llama-3.1-8bINT40.250.501.002.004.007.81
llama-3.3-70bFP160.440.881.753.507.0013.67
llama-3.3-70bINT40.220.440.881.753.506.84

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU. FP8 KV equals INT4 KV — both 1 byte per element.

Two readings: Llama 3.1 8B INT4 grows 6.5 → 14.8 GB from 4K to 128K — the 8 GB card that runs it at 16K (7.3) fails at 32K (8.4). Qwen2.5-Coder 32B FP16 goes 74.2 → 101.2: the gap between one 80 GB GPU (marginal at 4K) and needing 141 GB. Detail: Llama 3.1 8B VRAM.

The 70B-on-24GB question, answered

Llama 3.3 70B (70.6B parameters, 128K window). Weights alone: FP16 141.2 / FP8 70.6 / INT4 35.3 GB.

ContextFP16FP8/INT8INT4
4,096157.679.740.8
8,192158.079.941.1
16,384159.080.441.6
32,768160.981.342.5
65,536164.883.344.4
128,000172.186.948.1

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-07; batch 1, single GPU.

  • Llama 3.3 70B at INT4 needs 40.8 GB at 4K context — a single 24 GB GPU cannot hold it; the repo's GPU evaluator marks 24 GB Insufficient and computes 2 GPUs.
  • INT4 fits one 48 GB GPU up to 64K context (44.4 GB); at the full 128K window it reaches 48.1 GB — just over a single 48 GB card.
  • FP8 needs 79.7 GB at 4K: Marginal on an 80 GB H100/A100 (0.3 GB headroom), no fit from 16K (80.4 GB), comfortable on 96 GB+.
  • FP16 needs 157.6 GB: beyond H200's 141 GB, fits B200/MI300X 192 GB (headroom 34.4 GB).

Fit verdict only — deployment: how to run it locally; sizing: model hub and VRAM breakdown.

FAQ

Does Llama 3.3 70B run on a 24 GB GPU? No. INT4 needs 40.8 GB at 4K; the evaluator computes 2 GPUs. 48 GB covers INT4 up to 64K; 96 GB+ for FP8; 192 GB for FP16.

What is the difference between "fits" and "technically runs"? "Fits" means the full stack — weights + KV-cache + runtime + 10% margin — lives in one GPU's VRAM at batch 1 and 4K context. "Technically runs" usually means sharding or spilling to system RAM. This article only makes the first claim.

Can I offload to CPU or system RAM instead? Valid — but sharding and offload belong to the serving stack (local LLM serving engines); this engine doesn't model them, so no offload figures appear here.

My 8B model fits an 8 GB card at 4K. Why did it stop fitting? Llama 3.1 8B INT4: 6.5 GB at 4K, 7.3 at 16K, 8.4 at 32K — weights never changed; the KV-cache did.

Does any of this apply to Mac or unified memory? No — every figure assumes a single discrete GPU with dedicated VRAM; unified memory isn't modeled.

Limitations and assumptions

  • Memory fit only. No throughput, latency, bandwidth, quality, or price claims — pricing: cheapest GPUs.
  • Conditions: batch 1, single GPU, inference, default context 4,096; run 2026-10-07. Figures are CALCULATED_ESTIMATES — "needs"/"estimated", never measured. Labels: Fits (headroom ≥ 10% of the requirement), Marginal (< 10%), Does not fit.
  • Context windows are the registry's: gemma-2 models 8,000; phi-4 16,000; mistral-7b and codestral 32,000; mixtral-8x22b 65,000; most others 128,000. Runs never exceed 128K.
  • KV modeling: exact GQA (high confidence) for llama-3.1-8b, llama-3.3-70b, llama-3.1-70b, llama-3.1-405b, qwen-2.5-coder-32b, qwen-2.5-72b, mistral-nemo-12b, mistral-small-24b, deepseek-v3; modeled estimates (medium confidence) for everything else here — gemma, phi, 7B/14B class, distills, command-r, mistral-large, mixtral, ling. Both are engine output; only the confidence differs.
  • Not covered: CPU/RAM offload, multi-GPU throughput, Mac/unified memory, framework versions, quality effects, Q5/Q6 formats.

Related resources