Can I Run a 70B LLM on a 24GB GPU? Realistic Models, Quantization and Speed
Whether a 70B model runs on one 24GB GPU — Q4_K_M vs 2-bit GGUF sizes, CPU offload mechanics, tokens/sec bands, and when a 32B or multi-GPU setup is the better answer.
Direct answer
Yes — with an asterisk. A 70B model cannot run fully inside one 24 GB GPU at any reasonable 4-bit quality: the standard Q4_K_M GGUF for Llama 3.3 70B is a 42.5 GB file (bartowski, Hugging Face, verified 2026-10-11). What a single 24 GB card can do:
| What you want | On one 24 GB GPU | What it takes |
|---|---|---|
| FP16 / FP8 residency | No — 141 GB / 75 GB weights | Multiple GPUs or 48–80 GB class cards |
| Q4_K_M fully in VRAM | No — 42.5 GB file | 2× 24 GB or 1× 48 GB |
| Q4_K_M with CPU/RAM offload | Yes — GPU holds ~half the layers, system RAM holds the rest | ≥64 GB RAM; expect roughly 2–18 tok/s depending on split and memory bandwidth |
| 2-bit GGUF fully in VRAM (e.g. IQ2_XS, 21.1 GB) | Marginal yes — fits at short context, quality drops sharply | Short context (~4K); evaluate quality on your tasks |
| 70B-class quality, interactive speed | Usually no — a fully-resident 32B Q4 (~20 GB) is faster and often better value | See alternatives |
The four questions above are different questions. Conflating them is how people end up with a $1,600 card and a 2 tok/s chatbot.
What "run" actually means
- Load — can the weights be opened at all? Needs disk (42.5 GB free) and eventually a memory home.
- Fit fully in VRAM — weights + KV-cache + runtime all inside 24 GB. This is what "GPU-accelerated" normally means. For Llama 3.3 70B at Q4_K_M: no.
- Run with CPU offloading — llama.cpp/Ollama place as many transformer layers as fit on the GPU; the remainder execute on the CPU with weights in system RAM, crossing PCIe for every token. This works; speed depends on the split.
- Be practical for your use — interactive chat wants roughly ≥8–10 tok/s. Batch/offline work can tolerate less. A configuration that "runs" at 1.5 tok/s is technically correct and practically useless for chat.
Our VRAM tiers article already answers the fit question for every tier: Llama 3.3 70B at INT4 needs 40.8 GB at 4K context (deterministic engine, batch 1) and does not fit a single 24 GB GPU. This article answers the rest: how to run it anyway, at what speed, and when to stop trying.
Memory by quantization: theory vs the file you download
The naive math — 70.6B parameters × 0.5 bytes (INT4) = 35.3 GB — is a theoretical floor our calculator uses for its INT4 column. Real GGUF K-quants average closer to 4.8 bits per weight (mixed-precision blocks, some tensors kept higher-precision), so the on-disk file is larger. Always size from the actual file size, not from parameters × bytes.
Llama 3.3 70B Instruct GGUF sizes (bartowski repo, Hugging Face, verified 2026-10-11):
| Quant | File size | vs 24 GB VRAM | Notes |
|---|---|---|---|
| F16 | 141.1 GB | Impossible | Reference weights |
| Q8_0 | 75.0 GB | Impossible | Near-lossless; needs 80 GB+ class |
| Q5_K_M | 50.0 GB | Impossible | High quality; 48 GB cards marginal |
| Q4_K_M | 42.5 GB | Needs offload or 2×GPU | Community default for 70B |
| Q4_K_S | 40.4 GB | Needs offload | Slightly smaller, slightly lower quality |
| Q3_K_M | 34.3 GB | Needs offload | "Low quality" per repo label |
| IQ3_M | 31.9 GB | Needs offload | 3-bit I-quant |
| Q3_K_S | 30.9 GB | Needs offload | |
| IQ2_S | 22.2 GB | Fits at short context (marginal) | 2-bit — quality drops sharply |
| IQ2_XS | 21.1 GB | Fits at short context | Often the largest practical full-residency quant on 24 GB |
| IQ2_XXS | 19.1 GB | Fits | 2-bit, more degradation |
| Q2_K | 26.4 GB | Does not fit |
The runtime total is file size + KV-cache + overhead. KV-cache scales with context length (our KV-cache explainer covers the formula); budget roughly 1–2 GB at 4K context for a 70B GQA model at INT8 KV, more at 32K+. Runtime/CUDA overhead is ~1.5–2 GB. So:
Q4_K_M full residency ≈ 42.5 + ~1.5 (KV@4K) + ~1.6 (overhead) ≈ 45–46 GB
IQ2_XS full residency ≈ 21.1 + ~1.5 + ~1.6 ≈ 24.2 GB (marginal at 4K;
OOM risk at longer context)
That is why the fit tables say 48 GB is the first comfortable single-card tier for 70B INT4-class weights, and why a 24 GB card's only full-residency option is 2-bit — and 2-bit quality on a 70B model is a real compromise, not a free lunch. Evaluate on your own tasks before committing; we do not quote a single "quality loss %" because it is task-dependent (same policy as our DeepSeek R1 quantization note).
What works on a 24 GB GPU in practice
Three realistic configurations, with community-measured speed ranges (labeled approximate — see limitations):
1. Q4_K_M + CPU offload (best quality-per-VRAM)
- Split: ~20–22 GB of weights on GPU (~40–45 of 80 layers), ~20–23 GB in system RAM.
- System RAM: 64 GB is the comfortable target; 32 GB requires swap and gets ugly.
- Speed (community reports, mid-2026): tuned llama.cpp with most layers on GPU and fast DDR5: ~15–18 tok/s; Ollama's automatic split on a typical desktop: ~2–4 tok/s. The spread is the split ratio, CPU single-thread speed, and memory bandwidth — not model luck.
- Best for: quality-sensitive offline generation, RAG batch jobs, experimentation where waiting is acceptable.
2. 2-bit GGUF fully on GPU (IQ2_XS, ~21 GB)
- Residency: everything in VRAM at short context — no PCIe per-token penalty.
- Speed: roughly 15–25 tok/s reported via Ollama on a 4090-class card (community measurement, Jun 2026) — faster than a bad offload split because the whole model streams from GDDR.
- Quality: 2-bit is where 70B models start losing reasoning fidelity, long-context recall, and formatting discipline. Treat it as an experiment, not a deployment.
3. Don't: FP8/AWQ/GPTQ on one 24 GB card
vLLM, TensorRT-LLM and production AWQ/GPTQ stacks assume full GPU residency. 70B at INT8 is ~75 GB; there is no supported single-24GB path. CPU offloading is a llama.cpp/GGUF-family capability (also exposed through Ollama), not a vLLM feature.
GPU-only versus CPU/RAM offloading
Token generation step:
GPU-resident layers: weights stream from VRAM (~1 TB/s on a 4090)
CPU-resident layers: weights stream from RAM (~50–90 GB/s DDR5)
+ PCIe transfer per layer (~25–50 GB/s practical)
Every CPU-resident layer adds a memory-bandwidth and PCIe tax to every generated token. Prompt processing (prefill) batches matrix multiplies and fares better; decode is where offload hurts. Practical consequences:
- Doubling system RAM does not double speed — it just lets more layers exist off-GPU.
- Faster RAM (DDR5-6000+ dual/quad channel) helps more than more cores for decode.
-tthread count matters: too many threads oversubscribe and slow down; llama.cpp docs and community guides recommend matching physical cores and tuning down from there.- The only way to remove the tax is full residency — smaller quant on the same card, more cards, or a bigger card.
Engines and configuration
Verified commands and behaviors (checked 2026-10-11):
llama.cpp — explicit control, the reference offload implementation:
# Q4_K_M, offload as many layers as fit in 24 GB, 4K context, 8 threads
llama-server -m Llama-3.3-70B-Instruct-Q4_K_M.gguf -ngl 999 -c 4096 -t 8
-ngl 999 (alias --n-gpu-layers 999) means "offload everything that fits"; llama.cpp computes the split. Lower it (e.g. -ngl 40) to force more onto the CPU when experimenting. Verify the load log shows layers offloaded and the expected VRAM used. Pair with -fa (flash attention) to shrink KV-cache. Source: llama.cpp documentation and quantization README (ggml-org, 2026).
Ollama — automatic split, zero-config path:
ollama run llama3.3:70b # 43 GB Q4_K_M-class download; auto-splits GPU/CPU
ollama ps # check PROCESSOR column — "100% GPU" is NOT what you'll see
Ollama downloads the ~43 GB quantized library build and layers whatever fits. ollama ps reporting a partial GPU percentage confirms offload is active — if you see CPU where you expected GPU, that is your bottleneck (same diagnostic as our local AI setup guide). Ollama does not expose per-layer counts like -ngl; fine-tuning the split means llama.cpp directly.
ExLlamaV2 / AWQ / GPTQ / vLLM — GPU-only pipelines. Wrong tool for single-24GB 70B; right tools once you have 48 GB+ or multi-GPU. See local LLM serving engines for the architecture differences.
Better alternatives for interactive use
If the goal is "best model I can chat with on this hardware," 70B-on-24GB is usually the wrong trade:
| Alternative | Fits 24 GB fully? | Why it often wins |
|---|---|---|
| Qwen 2.5 32B / Qwen 2.5 Coder 32B at Q4 (~20 GB) | Yes — model page | Full GPU residency → true 4090-class speeds; 32B Q4 quality is competitive for most tasks |
| Llama 3.3 70B at 2-bit | Marginal | Full-residency speed, compromised quality — only if you specifically need 70B behaviors |
| 70B Q4_K_M on 2× RTX 3090/4090 | Yes (42.5 GB across two 24 GB cards) | Community-measured ~16–19 tok/s (llama.cpp tensor split; janreges benchmark dataset, Llama-3 70B Q4_K_M). See our 4× 4090 run page for the multi-card sizing pattern |
| Rent a 48 GB+ card for a few hours | L40S 48 GB: full Q4 residency, ~15 tok/s measured | Avoids capex; see cloud vs self-hosting for the breakeven frame and cheapest GPUs for current rates |
| Use an API | n/a | Zero hardware; pay per token — the default when local speed is a bottleneck |
The consumer GPU guide covers the buying decision; the one-line summary here is: 24 GB is the 32B-Q4 tier, not the 70B-Q4 tier.
When upgrading VRAM is worth it
| Jump | What unlocks | Rough used/new street logic |
|---|---|---|
| 24 GB → 2×24 GB | 70B Q4_K_M full residency at ~16–19 tok/s | Second used 3090/4090 often cheaper than one 48 GB workstation card |
| 24 GB → 48 GB (L40S, RTX 6000 Ada, used A6000) | 70B Q4 full residency single-card; 32B at FP8; headroom for long context | L40S page for rental/ownership context; measured ~15 tok/s Q4_K_M in the janreges set — more VRAM ≠ more speed; bandwidth and build differ |
| 48 GB → 80 GB (A100/H100) | 70B FP8; long-context KV headroom | Datacenter tier — see H100 vs A100 vs L40S for LLM inference and H100 vs H200 ROI analysis for that class |
Rule of thumb: upgrade VRAM when you keep hitting offload or OOM on models you actually use; don't upgrade hoping 70B-on-48GB will feel like a 14B-on-24GB — it won't; weights-per-token math still applies.
FAQ
Does ollama run llama3.3:70b work on a 4090 with 64 GB RAM?
Yes, commonly reported — as a CPU/GPU hybrid. Check ollama ps for the GPU percentage. Expect single-digit tok/s unless the split is favorable; measure before relying on it.
Is Q3 better than 2-bit for fitting 24 GB? Q3_K_M is 34.3 GB — still needs offload. Only the 2-bit IQ2_* class (19–22 GB) fits fully. So the choice is "Q4 offloaded (better quality, slower)" vs "2-bit resident (faster, worse quality)" — not Q3 vs Q2 on-card.
Does more system RAM make offload fast? It makes offload possible (fit the remaining layers without swap), not fast. Speed is bound by RAM bandwidth and PCIe, not capacity, once the layers are resident.
What about Mac unified memory? Apple Silicon pools RAM and VRAM; a 64–128 GB Mac can hold a 70B Q4 entirely in unified memory at respectable speeds — a different trade from PCIe offload. We do not publish Mac tok/s figures here; community measurements vary widely by chip and quant.
Multi-GPU on consumer cards — NVLink? RTX 4090 has no NVLink; 3090 has 2-way NVLink but llama.cpp tensor splitting works over PCIe either way. The janreges measurements above are PCIe tensor-split numbers.
Limitations
- GGUF file sizes are from the bartowski/Llama-3.3-70B-Instruct-GGUF Hugging Face repository, verified 2026-10-11. Other quants and repos vary by a few percent.
- Theoretical INT4 floor (35.3 GB) is our engine's parameters × 0.5 B/param figure; real K-quant files are larger. Both numbers are correct for their respective definitions — this article uses file sizes for practical sizing.
- Tokens/sec figures are community measurements, not OpenGPU Radar benchmarks: gingerlabs.ai (Jun 2026, RTX 4090 + Ollama/llama.cpp/ExLlama ranges), smeltcore.com (Q4_K_XL hybrid ~17.8 tok/s), and the janreges GPU-Benchmarks-on-LLM-Inference dataset (llama.cpp, Llama-3 70B Q4_K_M, multi-GPU full residency: 2×4090 ≈ 19.1 tok/s, L40S ≈ 15.3 tok/s). Hardware, driver, context, and build versions differ across sources; treat ranges as bands, not guarantees. Measure your own split with
llama-benchor a timed chat loop. - KV-cache totals cited as "~1–2 GB at 4K" are approximations for a 70B GQA model at INT8 KV; exact values depend on engine implementation (flash attention, paged allocation). Use the calculator for our engine's totals and the tiers article for published fit numbers at 4K/32K/128K.
- Quality at 2-bit and Q3 is not quantified here; perplexity-style averages hide task-specific regressions. Run your own eval prompts.
- Command syntax (
-ngl,-c,-t,-fa,ollama ps) verified against llama.cpp and Ollama documentation 2026-10-11; flags evolve — check--helpon your build. - Offload speed depends on your CPU and RAM more than this article can model. A 7950X with DDR5-6000 and a 13th-gen i5 with DDR4-3200 will not land in the same band.
GGUF sizes: bartowski/Llama-3.3-70B-Instruct-GGUF (Hugging Face), accessed 2026-10-11. llama.cpp quantization reference sizes: ggml-org/llama.cpp tools/quantize README. Multi-GPU tok/s: janreges/GPU-Benchmarks-on-LLM-Inference (llama.cpp, Llama-3 70B Q4_K_M). Single-GPU offload ranges: gingerlabs.ai (2026-06), smeltcore.com. VRAM fit totals: OpenGPU Radar deterministic engine via VRAM tiers, run 2026-10-07.