How to Fine-Tune an LLM with Limited GPU Memory in 2026
Fine-tune an LLM on limited GPU memory: full fine-tune vs LoRA vs QLoRA memory math, fit tables for 12-80GB GPUs, and checkpointing tricks.
Direct answer
Full fine-tuning costs 16 bytes per parameter of training state — Llama 3.1 8B needs 128.5 GB (past every 80 GB card), Qwen 2.5 Coder 32B needs 520.0 GB, Llama 3.3 70B needs 1,129.6 GB (past any single GPU). With limited VRAM the question is not whether to use an adapter, but which: LoRA freezes the base at its inference footprint and trains roughly 0.2% of the parameters, and QLoRA drops the frozen base to 4 bits, making a 48 GB card the floor for 70B fine-tuning. The arithmetic, run 2026-10-09:
| Method | Bytes per param to train | Llama 3.1 8B | Qwen 2.5 Coder 32B | Llama 3.3 70B |
|---|---|---|---|---|
| Full fine-tune (AdamW, mixed precision) | 16 | 128.5 GB | 520.0 GB | 1,129.6 GB |
| LoRA, FP16 base frozen | 2 + adapter | 16.1 GB | 65.0 GB | 141.2 GB |
| QLoRA, INT4 base frozen | 0.5 + adapter | 4.0 GB | 16.2 GB | 35.3 GB |
Source: arithmetic — OpenGPU Radar model registry parameter counts × bytes/param (FP16 2.0, FP8/INT8 1.0, INT4 0.5, full AdamW stack 16); run 2026-10-09. LoRA/QLoRA rows are the frozen base weights — adapter state adds well under 1 GB at these ranks. Activations excluded (covered below).
Read down the columns: on a 24 GB card the full-fine-tune column never clears 1.3B parameters, while QLoRA reaches the 42B class — and at 48 GB, 70B QLoRA (35.3 GB of base weights plus adapter, optimizer and activation budget) clears with room. That is the whole game on limited hardware.
Why full fine-tuning blows past one GPU
Mixed-precision AdamW keeps five copies of state per parameter:
| State | Bytes/param | Llama 3.1 8B |
|---|---|---|
| FP16 weights | 2 | 16.1 GB |
| FP16 gradients | 2 | 16.1 GB |
| FP32 master weights | 4 | 32.1 GB |
| FP32 momentum | 4 | 32.1 GB |
| FP32 variance | 4 | 32.1 GB |
| Total | 16 | 128.5 GB |
Source: arithmetic (parameter count × bytes/param), run 2026-10-09. Conditions: AdamW mixed-precision training bookkeeping; activations excluded.
The optimizer side — master weights plus momentum plus variance, 12 bytes — is three times the 2-byte weight footprint itself, and with gradients the non-weight state reaches 14 bytes per parameter against 2 for inference. This is exactly the "optimizer state beats weights" effect the site's fine-tuning economics page uses to screen candidates: its methodology excludes full FP16 fine-tuning of a 70B (≈140 GB of weights alone) from single-GPU candidates by design.
Multi-GPU arithmetic for the 70B case: 1,129.6 GB of state against one 8×H100 node's 640 GB does not fit — sharding divides state, it doesn't add capacity. Two nodes — 16×80 GB = 1,280 GB — leave 150.4 GB (13%) of aggregate headroom for activations and transient gathers. On the single-GPU side, the 8B case (128.5 GB) clears only 141 GB-and-up cards — an H200 leaves 12.5 GB for activations after the state, and 192 GB-class parts are where the margin becomes real. Either way it is a datacenter budget, not a consumer one. The practical full-fine-tune ceiling on one 80 GB GPU is the 4.8B class.
LoRA: freeze the base, train two small matrices
LoRA (Low-Rank Adaptation) leaves the base weights untouched and trains a pair of low-rank matrices per adapted linear layer. The trainable count is arithmetic: r × (d_in + d_out) per adapted matrix, where r is the adapter rank.
Worked example — Llama 3.1 8B, rank 16, attention projections q/k/v/o in every layer (registry architecture: 32 layers, 32 attention heads × 128 head-dim = 4,096 hidden):
- Per matrix: 16 × (4,096 + 4,096) = 131,072 parameters
- Per layer × 4 matrices: 524,288; × 32 layers: 16,777,216 parameters ≈ 0.21% of 8.03B
- Adapter + gradients + optimizer at the same 16 bytes/param: 0.27 GB
Source: arithmetic from OpenGPU Radar registry architecture fields (layers, heads, head-dim), run 2026-10-09.
So the residency bill for 8B LoRA is the frozen FP16 base (16.1 GB) plus that 0.27 GB plus activations — comfortably inside 24 GB, and the base's inference-side total from the VRAM engine at 4K context is 20.0 GB (weights 16.1 + KV 0.5 + runtime 1.6 + 10% margin), consistent with a single RTX 4090.
What LoRA buys by model class, single GPU with a 3 GB adapter/optimizer/activation budget on top:
- 8B class on 24 GB: FP16 base (16.1 GB) fits with headroom.
- 32B class on 48 GB: FP16 base (65.0 GB) does not fit; an INT8/FP8 base (32.5 GB) does — which is why the fine-tuning economics page lists Qwen 2.5 Coder 32B as an "FP16/INT8 base" LoRA target across 80 GB-class candidates and 48 GB cards alike.
- 70B class on 80 GB: FP16 base (141.2 GB) never fits one card; INT8 base (70.6 GB) fits; INT4 base is QLoRA (next).
QLoRA: the 4-bit base
QLoRA keeps the base frozen at INT4 — 0.5 bytes per parameter — and trains the adapter with the same bookkeeping as LoRA. The base numbers are the weights figures the site publishes everywhere else: 70B = 35.3 GB, 32B = 16.25 GB, 8B = 4.0 GB.
- 70B QLoRA: 35.3 GB base + adapter/optimizer/activation budget ≈ 38 GB — inside 48 GB. The fine-tuning economics page's methodology sets its fit test at "48 GB-class cards and up" for exactly this reason, and the VRAM tiers article puts the 70B INT4 inference-side total at 40.8 GB at 4K — both figures agree that one L40S-class card holds it.
- 32B QLoRA: 16.25 GB base fits a 24 GB RTX 4090 with budget to spare.
- Quality: the frozen base carries INT4's known trade-off — the quantization explainer puts INT4/AWQ at a 3–5% perplexity cost. The adapter itself trains in higher precision; the base's quantization is what you inherit. Benchmark on your own data before shipping — that is the same advice the quantization guide gives for inference.
Activations, batch, and sequence: what arithmetic can't fix
Everything above is resident state. Activations — the forward-pass tensors kept for the backward pass — scale linearly with batch size × sequence length and are the reason fit tables state their budgets explicitly rather than pretending one number is final:
- Halving batch size halves activation memory. So does halving sequence length — at a cost in training quality and throughput that you should only pay when the card forces it.
- Gradient checkpointing stores activations at checkpoint boundaries instead of every layer: the stored footprint shrinks by the number of layers per segment, and the backward pass recomputes the skipped segments. The trade is compute for memory, in percentages we don't publish because we haven't measured them.
- Training precision: BF16/FP16 remains the training baseline (quantization explainer); FP8 training exists on Hopper/Blackwell-class parts but is not what a limited-VRAM setup is optimizing for.
Single-GPU ceilings: the fit table
(VRAM − 3 GB) ÷ bytes/param, where the 3 GB is CUDA context + adapter states + a modest batch-1 activation allowance:
| GPU | VRAM | Full fine-tune (÷16) | LoRA FP16 base (÷2) | LoRA INT8 base (÷1) | QLoRA INT4 base (÷0.5) |
|---|---|---|---|---|---|
| RTX 3060 | 12 GB | ≤ 0.6B | ≤ 4.5B | ≤ 9B | ≤ 18B |
| T4 | 16 GB | ≤ 0.8B | ≤ 6.5B | ≤ 13B | ≤ 26B |
| RTX 4090 | 24 GB | ≤ 1.3B | ≤ 10.5B | ≤ 21B | ≤ 42B |
| L40S | 48 GB | ≤ 2.8B | ≤ 22.5B | ≤ 45B | ≤ 90B |
| A100 80GB / H100 | 80 GB | ≤ 4.8B | ≤ 38.5B | ≤ 77B | ≤ 154B |
Source: arithmetic — (VRAM − 3 GB budget) ÷ bytes/param (16 / 2 / 1 / 0.5), run 2026-10-09. These are arithmetic ceilings, not measured training runs: real activations at larger batch × sequence shrink every column, the full-fine-tune column most of all. Inference-side fit totals (weights + KV + runtime + 10% margin) come from the VRAM calculator.
How to read it: the full-fine-tune column is a hard upper bound (activations make it worse, never better); the adapter columns hold up in practice for batch-1-to-small runs and degrade gracefully — drop the batch, checkpoint, and you stay in the cell.
What to actually rent
- Rate arithmetic belongs on the fine-tuning economics page — it tracks observed provider rows for LoRA/QLoRA candidates (A100, H100, L40S, H200) and its cost driver applies here unchanged: run-hours × hourly rate, because a 20% cheaper GPU rarely compensates for a 2× slower iteration loop.
- For side-by-side rates while you decide: GPU comparison and cheapest GPUs.
- The budget-shaped version of this decision — renting capability rather than reflex — appears as an 8B fine-tuning example in the automated GPU comparison article.
FAQ
Can I fine-tune an 8B model on a 24GB GPU? LoRA yes: FP16 base 16.1 GB + adapter + activations fits (inference-side reference total at 4K: 20.0 GB from the engine). Full fine-tuning no: 128.5 GB of training state — five times the card.
Is QLoRA worse than LoRA? Different frozen base precision, identical adapter training. You inherit INT4's 3–5% perplexity trade-off from the quantization explainer; whether that matters is a benchmark question about your data, not a rule.
Does gradient checkpointing make full fine-tuning fit on one GPU? No. Checkpointing shrinks the activation term; it does not touch the 16-bytes-per-parameter optimizer stack. An 8B full fine-tune is still 128.5 GB before activations.
What fits a rented 80GB card? Arithmetic ceilings: full fine-tune ≤ 4.8B, LoRA to 38.5B (FP16) or 77B (INT8), QLoRA to 154B — with activations eating into each. Verify inference-side base fits on the calculator.
Limitations and assumptions
- Training figures are arithmetic, not measurements. Every total = parameter count × bytes/param with the formulas shown above; all are CALCULATED_ESTIMATES (confidence: arithmetic), run 2026-10-09. No training throughput, convergence, or loss numbers appear in this article — we have not measured them, and this site does not publish benchmark claims it can't source.
- Conditions: AdamW mixed precision; 3 GB runtime/adapter/activation budget for ceilings (CUDA context + adapter + modest batch-1 activations); activations excluded from the Direct-answer table and scale linearly with batch × sequence.
- Engine figures (20.0 GB for 8B FP16 at 4K, 40.8 GB for 70B INT4, 74.2 GB for 32B FP16) are inference-side totals from the VRAM canonical engine — weights + KV-cache + 1.2 GB runtime + 0.4 GB activation + 10% fragmentation margin, batch 1, single GPU — the same code path as the VRAM tiers article.
- Not covered: multi-GPU sharding libraries, CPU/RAM offload, dataset-pipeline memory, LoRA-variant extras (QLoRA's paged optimizers, DoRA), and FP8 training. The single multi-GPU figure (16×80 GB for 70B full fine-tune) is aggregate arithmetic, not a deployment recipe.
Related resources
- LLM fine-tuning GPU economics — verified rates for LoRA/QLoRA candidates
- VRAM Calculator · What is quantization? · What is VRAM?
- Which LLMs fit 8–80GB GPUs — the inference-side companion tables
- Models: Llama 3.1 8B · Qwen 2.5 Coder 32B · Llama 3.3 70B
- GPUs: RTX 4090 · L40S · A100 80GB · H100 · H200
- Automated GPU comparison — 8B fine-tuning on a budget, worked side-by-side
Related Articles
Which LLMs Can Run on 8GB, 16GB, 24GB, 48GB and 80GB GPUs in 2026?
Which LLMs fit on 8GB, 16GB, 24GB, 48GB and 80GB GPUs — verified fit tables by quantization and context from OpenGPU Radar's VRAM engine.
InfrastructureB200 NVLink 5.0 Scaling: When Does 8x B200 Beat 16x H100?
8x B200 SXM (NVLink 5.0 at 1.8 TB/s) equals 16x H100's aggregate NVLink bandwidth at similar cost — and delivers 50% more 70B inference throughput. Here's the math on bandwidth, power, and cost.
InfrastructureLLM VRAM Calculator: GPU Requirements for Llama, Qwen, DeepSeek & Claude in 2026
Find the right GPU for any LLM: Llama 3.3 70B needs 71 GB of FP8 weights (~81 GB in service), DeepSeek R1 671B needs 786 GB at 128K context. VRAM sizing chart by model.
InfrastructureVerified Free LLM APIs: Which Providers Actually Require No Credit Card in 2026?
Ranked comparison of free LLM APIs requiring no credit card: Google AI Studio (Gemini 2.0 Flash), Groq (Llama 3.3 70B), Cerebras, Cloudflare Workers AI, and more. Verified 2026-09-26.