Local AI2026-10-11•By Sree•5 min read

Can I Run a 70B LLM on a 24GB GPU? Realistic Models, Quantization and Speed

Whether a 70B model runs on one 24GB GPU — Q4_K_M vs 2-bit GGUF sizes, CPU offload mechanics, tokens/sec bands, and when a 32B or multi-GPU setup is the better answer.

Direct answer

Yes — with an asterisk. A 70B model cannot run fully inside one 24 GB GPU at any reasonable 4-bit quality: the standard Q4_K_M GGUF for Llama 3.3 70B is a 42.5 GB file (bartowski, Hugging Face, verified 2026-10-11). What a single 24 GB card can do:

What you wantOn one 24 GB GPUWhat it takes
FP16 / FP8 residencyNo — 141 GB / 75 GB weightsMultiple GPUs or 48–80 GB class cards
Q4_K_M fully in VRAMNo — 42.5 GB file2× 24 GB or 1× 48 GB
Q4_K_M with CPU/RAM offloadYes — GPU holds ~half the layers, system RAM holds the rest≥64 GB RAM; expect roughly 2–18 tok/s depending on split and memory bandwidth
2-bit GGUF fully in VRAM (e.g. IQ2_XS, 21.1 GB)Marginal yes — fits at short context, quality drops sharplyShort context (~4K); evaluate quality on your tasks
70B-class quality, interactive speedUsually no — a fully-resident 32B Q4 (~20 GB) is faster and often better valueSee alternatives

The four questions above are different questions. Conflating them is how people end up with a $1,600 card and a 2 tok/s chatbot.

What "run" actually means

  1. Load — can the weights be opened at all? Needs disk (42.5 GB free) and eventually a memory home.
  2. Fit fully in VRAM — weights + KV-cache + runtime all inside 24 GB. This is what "GPU-accelerated" normally means. For Llama 3.3 70B at Q4_K_M: no.
  3. Run with CPU offloading — llama.cpp/Ollama place as many transformer layers as fit on the GPU; the remainder execute on the CPU with weights in system RAM, crossing PCIe for every token. This works; speed depends on the split.
  4. Be practical for your use — interactive chat wants roughly ≥8–10 tok/s. Batch/offline work can tolerate less. A configuration that "runs" at 1.5 tok/s is technically correct and practically useless for chat.

Our VRAM tiers article already answers the fit question for every tier: Llama 3.3 70B at INT4 needs 40.8 GB at 4K context (deterministic engine, batch 1) and does not fit a single 24 GB GPU. This article answers the rest: how to run it anyway, at what speed, and when to stop trying.

Memory by quantization: theory vs the file you download

The naive math — 70.6B parameters × 0.5 bytes (INT4) = 35.3 GB — is a theoretical floor our calculator uses for its INT4 column. Real GGUF K-quants average closer to 4.8 bits per weight (mixed-precision blocks, some tensors kept higher-precision), so the on-disk file is larger. Always size from the actual file size, not from parameters × bytes.

Llama 3.3 70B Instruct GGUF sizes (bartowski repo, Hugging Face, verified 2026-10-11):

QuantFile sizevs 24 GB VRAMNotes
F16141.1 GBImpossibleReference weights
Q8_075.0 GBImpossibleNear-lossless; needs 80 GB+ class
Q5_K_M50.0 GBImpossibleHigh quality; 48 GB cards marginal
Q4_K_M42.5 GBNeeds offload or 2×GPUCommunity default for 70B
Q4_K_S40.4 GBNeeds offloadSlightly smaller, slightly lower quality
Q3_K_M34.3 GBNeeds offload"Low quality" per repo label
IQ3_M31.9 GBNeeds offload3-bit I-quant
Q3_K_S30.9 GBNeeds offload
IQ2_S22.2 GBFits at short context (marginal)2-bit — quality drops sharply
IQ2_XS21.1 GBFits at short contextOften the largest practical full-residency quant on 24 GB
IQ2_XXS19.1 GBFits2-bit, more degradation
Q2_K26.4 GBDoes not fit

The runtime total is file size + KV-cache + overhead. KV-cache scales with context length (our KV-cache explainer covers the formula); budget roughly 1–2 GB at 4K context for a 70B GQA model at INT8 KV, more at 32K+. Runtime/CUDA overhead is ~1.5–2 GB. So:

Q4_K_M full residency ≈ 42.5 + ~1.5 (KV@4K) + ~1.6 (overhead) ≈ 45–46 GB
IQ2_XS full residency  ≈ 21.1 + ~1.5 + ~1.6               ≈ 24.2 GB  (marginal at 4K;
                                                                       OOM risk at longer context)

That is why the fit tables say 48 GB is the first comfortable single-card tier for 70B INT4-class weights, and why a 24 GB card's only full-residency option is 2-bit — and 2-bit quality on a 70B model is a real compromise, not a free lunch. Evaluate on your own tasks before committing; we do not quote a single "quality loss %" because it is task-dependent (same policy as our DeepSeek R1 quantization note).

What works on a 24 GB GPU in practice

Three realistic configurations, with community-measured speed ranges (labeled approximate — see limitations):

1. Q4_K_M + CPU offload (best quality-per-VRAM)

  • Split: ~20–22 GB of weights on GPU (~40–45 of 80 layers), ~20–23 GB in system RAM.
  • System RAM: 64 GB is the comfortable target; 32 GB requires swap and gets ugly.
  • Speed (community reports, mid-2026): tuned llama.cpp with most layers on GPU and fast DDR5: ~15–18 tok/s; Ollama's automatic split on a typical desktop: ~2–4 tok/s. The spread is the split ratio, CPU single-thread speed, and memory bandwidth — not model luck.
  • Best for: quality-sensitive offline generation, RAG batch jobs, experimentation where waiting is acceptable.

2. 2-bit GGUF fully on GPU (IQ2_XS, ~21 GB)

  • Residency: everything in VRAM at short context — no PCIe per-token penalty.
  • Speed: roughly 15–25 tok/s reported via Ollama on a 4090-class card (community measurement, Jun 2026) — faster than a bad offload split because the whole model streams from GDDR.
  • Quality: 2-bit is where 70B models start losing reasoning fidelity, long-context recall, and formatting discipline. Treat it as an experiment, not a deployment.

3. Don't: FP8/AWQ/GPTQ on one 24 GB card

vLLM, TensorRT-LLM and production AWQ/GPTQ stacks assume full GPU residency. 70B at INT8 is ~75 GB; there is no supported single-24GB path. CPU offloading is a llama.cpp/GGUF-family capability (also exposed through Ollama), not a vLLM feature.

GPU-only versus CPU/RAM offloading

Token generation step:
  GPU-resident layers:  weights stream from VRAM (~1 TB/s on a 4090)
  CPU-resident layers:  weights stream from RAM  (~50–90 GB/s DDR5)
                        + PCIe transfer per layer (~25–50 GB/s practical)

Every CPU-resident layer adds a memory-bandwidth and PCIe tax to every generated token. Prompt processing (prefill) batches matrix multiplies and fares better; decode is where offload hurts. Practical consequences:

  • Doubling system RAM does not double speed — it just lets more layers exist off-GPU.
  • Faster RAM (DDR5-6000+ dual/quad channel) helps more than more cores for decode.
  • -t thread count matters: too many threads oversubscribe and slow down; llama.cpp docs and community guides recommend matching physical cores and tuning down from there.
  • The only way to remove the tax is full residency — smaller quant on the same card, more cards, or a bigger card.

Engines and configuration

Verified commands and behaviors (checked 2026-10-11):

llama.cpp — explicit control, the reference offload implementation:

# Q4_K_M, offload as many layers as fit in 24 GB, 4K context, 8 threads
llama-server -m Llama-3.3-70B-Instruct-Q4_K_M.gguf -ngl 999 -c 4096 -t 8

-ngl 999 (alias --n-gpu-layers 999) means "offload everything that fits"; llama.cpp computes the split. Lower it (e.g. -ngl 40) to force more onto the CPU when experimenting. Verify the load log shows layers offloaded and the expected VRAM used. Pair with -fa (flash attention) to shrink KV-cache. Source: llama.cpp documentation and quantization README (ggml-org, 2026).

Ollama — automatic split, zero-config path:

ollama run llama3.3:70b     # 43 GB Q4_K_M-class download; auto-splits GPU/CPU
ollama ps                   # check PROCESSOR column — "100% GPU" is NOT what you'll see

Ollama downloads the ~43 GB quantized library build and layers whatever fits. ollama ps reporting a partial GPU percentage confirms offload is active — if you see CPU where you expected GPU, that is your bottleneck (same diagnostic as our local AI setup guide). Ollama does not expose per-layer counts like -ngl; fine-tuning the split means llama.cpp directly.

ExLlamaV2 / AWQ / GPTQ / vLLM — GPU-only pipelines. Wrong tool for single-24GB 70B; right tools once you have 48 GB+ or multi-GPU. See local LLM serving engines for the architecture differences.

Better alternatives for interactive use

If the goal is "best model I can chat with on this hardware," 70B-on-24GB is usually the wrong trade:

AlternativeFits 24 GB fully?Why it often wins
Qwen 2.5 32B / Qwen 2.5 Coder 32B at Q4 (~20 GB)Yes — model pageFull GPU residency → true 4090-class speeds; 32B Q4 quality is competitive for most tasks
Llama 3.3 70B at 2-bitMarginalFull-residency speed, compromised quality — only if you specifically need 70B behaviors
70B Q4_K_M on 2× RTX 3090/4090Yes (42.5 GB across two 24 GB cards)Community-measured ~16–19 tok/s (llama.cpp tensor split; janreges benchmark dataset, Llama-3 70B Q4_K_M). See our 4× 4090 run page for the multi-card sizing pattern
Rent a 48 GB+ card for a few hoursL40S 48 GB: full Q4 residency, ~15 tok/s measuredAvoids capex; see cloud vs self-hosting for the breakeven frame and cheapest GPUs for current rates
Use an APIn/aZero hardware; pay per token — the default when local speed is a bottleneck

The consumer GPU guide covers the buying decision; the one-line summary here is: 24 GB is the 32B-Q4 tier, not the 70B-Q4 tier.

When upgrading VRAM is worth it

JumpWhat unlocksRough used/new street logic
24 GB → 2×24 GB70B Q4_K_M full residency at ~16–19 tok/sSecond used 3090/4090 often cheaper than one 48 GB workstation card
24 GB → 48 GB (L40S, RTX 6000 Ada, used A6000)70B Q4 full residency single-card; 32B at FP8; headroom for long contextL40S page for rental/ownership context; measured ~15 tok/s Q4_K_M in the janreges set — more VRAM ≠ more speed; bandwidth and build differ
48 GB → 80 GB (A100/H100)70B FP8; long-context KV headroomDatacenter tier — see H100 vs A100 vs L40S for LLM inference and H100 vs H200 ROI analysis for that class

Rule of thumb: upgrade VRAM when you keep hitting offload or OOM on models you actually use; don't upgrade hoping 70B-on-48GB will feel like a 14B-on-24GB — it won't; weights-per-token math still applies.

FAQ

Does ollama run llama3.3:70b work on a 4090 with 64 GB RAM? Yes, commonly reported — as a CPU/GPU hybrid. Check ollama ps for the GPU percentage. Expect single-digit tok/s unless the split is favorable; measure before relying on it.

Is Q3 better than 2-bit for fitting 24 GB? Q3_K_M is 34.3 GB — still needs offload. Only the 2-bit IQ2_* class (19–22 GB) fits fully. So the choice is "Q4 offloaded (better quality, slower)" vs "2-bit resident (faster, worse quality)" — not Q3 vs Q2 on-card.

Does more system RAM make offload fast? It makes offload possible (fit the remaining layers without swap), not fast. Speed is bound by RAM bandwidth and PCIe, not capacity, once the layers are resident.

What about Mac unified memory? Apple Silicon pools RAM and VRAM; a 64–128 GB Mac can hold a 70B Q4 entirely in unified memory at respectable speeds — a different trade from PCIe offload. We do not publish Mac tok/s figures here; community measurements vary widely by chip and quant.

Multi-GPU on consumer cards — NVLink? RTX 4090 has no NVLink; 3090 has 2-way NVLink but llama.cpp tensor splitting works over PCIe either way. The janreges measurements above are PCIe tensor-split numbers.

Limitations

  • GGUF file sizes are from the bartowski/Llama-3.3-70B-Instruct-GGUF Hugging Face repository, verified 2026-10-11. Other quants and repos vary by a few percent.
  • Theoretical INT4 floor (35.3 GB) is our engine's parameters × 0.5 B/param figure; real K-quant files are larger. Both numbers are correct for their respective definitions — this article uses file sizes for practical sizing.
  • Tokens/sec figures are community measurements, not OpenGPU Radar benchmarks: gingerlabs.ai (Jun 2026, RTX 4090 + Ollama/llama.cpp/ExLlama ranges), smeltcore.com (Q4_K_XL hybrid ~17.8 tok/s), and the janreges GPU-Benchmarks-on-LLM-Inference dataset (llama.cpp, Llama-3 70B Q4_K_M, multi-GPU full residency: 2×4090 ≈ 19.1 tok/s, L40S ≈ 15.3 tok/s). Hardware, driver, context, and build versions differ across sources; treat ranges as bands, not guarantees. Measure your own split with llama-bench or a timed chat loop.
  • KV-cache totals cited as "~1–2 GB at 4K" are approximations for a 70B GQA model at INT8 KV; exact values depend on engine implementation (flash attention, paged allocation). Use the calculator for our engine's totals and the tiers article for published fit numbers at 4K/32K/128K.
  • Quality at 2-bit and Q3 is not quantified here; perplexity-style averages hide task-specific regressions. Run your own eval prompts.
  • Command syntax (-ngl, -c, -t, -fa, ollama ps) verified against llama.cpp and Ollama documentation 2026-10-11; flags evolve — check --help on your build.
  • Offload speed depends on your CPU and RAM more than this article can model. A 7950X with DDR5-6000 and a 13th-gen i5 with DDR4-3200 will not land in the same band.

GGUF sizes: bartowski/Llama-3.3-70B-Instruct-GGUF (Hugging Face), accessed 2026-10-11. llama.cpp quantization reference sizes: ggml-org/llama.cpp tools/quantize README. Multi-GPU tok/s: janreges/GPU-Benchmarks-on-LLM-Inference (llama.cpp, Llama-3 70B Q4_K_M). Single-GPU offload ranges: gingerlabs.ai (2026-06), smeltcore.com. VRAM fit totals: OpenGPU Radar deterministic engine via VRAM tiers, run 2026-10-07.