Infrastructure2026-10-09•By Sree•5 min read

How to Fine-Tune an LLM with Limited GPU Memory in 2026

Fine-tune an LLM on limited GPU memory: full fine-tune vs LoRA vs QLoRA memory math, fit tables for 12-80GB GPUs, and checkpointing tricks.

Direct answer

Full fine-tuning costs 16 bytes per parameter of training state — Llama 3.1 8B needs 128.5 GB (past every 80 GB card), Qwen 2.5 Coder 32B needs 520.0 GB, Llama 3.3 70B needs 1,129.6 GB (past any single GPU). With limited VRAM the question is not whether to use an adapter, but which: LoRA freezes the base at its inference footprint and trains roughly 0.2% of the parameters, and QLoRA drops the frozen base to 4 bits, making a 48 GB card the floor for 70B fine-tuning. The arithmetic, run 2026-10-09:

MethodBytes per param to trainLlama 3.1 8BQwen 2.5 Coder 32BLlama 3.3 70B
Full fine-tune (AdamW, mixed precision)16128.5 GB520.0 GB1,129.6 GB
LoRA, FP16 base frozen2 + adapter16.1 GB65.0 GB141.2 GB
QLoRA, INT4 base frozen0.5 + adapter4.0 GB16.2 GB35.3 GB

Source: arithmetic — OpenGPU Radar model registry parameter counts × bytes/param (FP16 2.0, FP8/INT8 1.0, INT4 0.5, full AdamW stack 16); run 2026-10-09. LoRA/QLoRA rows are the frozen base weights — adapter state adds well under 1 GB at these ranks. Activations excluded (covered below).

Read down the columns: on a 24 GB card the full-fine-tune column never clears 1.3B parameters, while QLoRA reaches the 42B class — and at 48 GB, 70B QLoRA (35.3 GB of base weights plus adapter, optimizer and activation budget) clears with room. That is the whole game on limited hardware.

Why full fine-tuning blows past one GPU

Mixed-precision AdamW keeps five copies of state per parameter:

StateBytes/paramLlama 3.1 8B
FP16 weights216.1 GB
FP16 gradients216.1 GB
FP32 master weights432.1 GB
FP32 momentum432.1 GB
FP32 variance432.1 GB
Total16128.5 GB

Source: arithmetic (parameter count × bytes/param), run 2026-10-09. Conditions: AdamW mixed-precision training bookkeeping; activations excluded.

The optimizer side — master weights plus momentum plus variance, 12 bytes — is three times the 2-byte weight footprint itself, and with gradients the non-weight state reaches 14 bytes per parameter against 2 for inference. This is exactly the "optimizer state beats weights" effect the site's fine-tuning economics page uses to screen candidates: its methodology excludes full FP16 fine-tuning of a 70B (≈140 GB of weights alone) from single-GPU candidates by design.

Multi-GPU arithmetic for the 70B case: 1,129.6 GB of state against one 8×H100 node's 640 GB does not fit — sharding divides state, it doesn't add capacity. Two nodes — 16×80 GB = 1,280 GB — leave 150.4 GB (13%) of aggregate headroom for activations and transient gathers. On the single-GPU side, the 8B case (128.5 GB) clears only 141 GB-and-up cards — an H200 leaves 12.5 GB for activations after the state, and 192 GB-class parts are where the margin becomes real. Either way it is a datacenter budget, not a consumer one. The practical full-fine-tune ceiling on one 80 GB GPU is the 4.8B class.

LoRA: freeze the base, train two small matrices

LoRA (Low-Rank Adaptation) leaves the base weights untouched and trains a pair of low-rank matrices per adapted linear layer. The trainable count is arithmetic: r × (d_in + d_out) per adapted matrix, where r is the adapter rank.

Worked example — Llama 3.1 8B, rank 16, attention projections q/k/v/o in every layer (registry architecture: 32 layers, 32 attention heads × 128 head-dim = 4,096 hidden):

  • Per matrix: 16 × (4,096 + 4,096) = 131,072 parameters
  • Per layer × 4 matrices: 524,288; × 32 layers: 16,777,216 parameters ≈ 0.21% of 8.03B
  • Adapter + gradients + optimizer at the same 16 bytes/param: 0.27 GB

Source: arithmetic from OpenGPU Radar registry architecture fields (layers, heads, head-dim), run 2026-10-09.

So the residency bill for 8B LoRA is the frozen FP16 base (16.1 GB) plus that 0.27 GB plus activations — comfortably inside 24 GB, and the base's inference-side total from the VRAM engine at 4K context is 20.0 GB (weights 16.1 + KV 0.5 + runtime 1.6 + 10% margin), consistent with a single RTX 4090.

What LoRA buys by model class, single GPU with a 3 GB adapter/optimizer/activation budget on top:

  • 8B class on 24 GB: FP16 base (16.1 GB) fits with headroom.
  • 32B class on 48 GB: FP16 base (65.0 GB) does not fit; an INT8/FP8 base (32.5 GB) does — which is why the fine-tuning economics page lists Qwen 2.5 Coder 32B as an "FP16/INT8 base" LoRA target across 80 GB-class candidates and 48 GB cards alike.
  • 70B class on 80 GB: FP16 base (141.2 GB) never fits one card; INT8 base (70.6 GB) fits; INT4 base is QLoRA (next).

QLoRA: the 4-bit base

QLoRA keeps the base frozen at INT4 — 0.5 bytes per parameter — and trains the adapter with the same bookkeeping as LoRA. The base numbers are the weights figures the site publishes everywhere else: 70B = 35.3 GB, 32B = 16.25 GB, 8B = 4.0 GB.

  • 70B QLoRA: 35.3 GB base + adapter/optimizer/activation budget ≈ 38 GB — inside 48 GB. The fine-tuning economics page's methodology sets its fit test at "48 GB-class cards and up" for exactly this reason, and the VRAM tiers article puts the 70B INT4 inference-side total at 40.8 GB at 4K — both figures agree that one L40S-class card holds it.
  • 32B QLoRA: 16.25 GB base fits a 24 GB RTX 4090 with budget to spare.
  • Quality: the frozen base carries INT4's known trade-off — the quantization explainer puts INT4/AWQ at a 3–5% perplexity cost. The adapter itself trains in higher precision; the base's quantization is what you inherit. Benchmark on your own data before shipping — that is the same advice the quantization guide gives for inference.

Activations, batch, and sequence: what arithmetic can't fix

Everything above is resident state. Activations — the forward-pass tensors kept for the backward pass — scale linearly with batch size × sequence length and are the reason fit tables state their budgets explicitly rather than pretending one number is final:

  • Halving batch size halves activation memory. So does halving sequence length — at a cost in training quality and throughput that you should only pay when the card forces it.
  • Gradient checkpointing stores activations at checkpoint boundaries instead of every layer: the stored footprint shrinks by the number of layers per segment, and the backward pass recomputes the skipped segments. The trade is compute for memory, in percentages we don't publish because we haven't measured them.
  • Training precision: BF16/FP16 remains the training baseline (quantization explainer); FP8 training exists on Hopper/Blackwell-class parts but is not what a limited-VRAM setup is optimizing for.

Single-GPU ceilings: the fit table

(VRAM − 3 GB) ÷ bytes/param, where the 3 GB is CUDA context + adapter states + a modest batch-1 activation allowance:

GPUVRAMFull fine-tune (÷16)LoRA FP16 base (÷2)LoRA INT8 base (÷1)QLoRA INT4 base (÷0.5)
RTX 306012 GB≤ 0.6B≤ 4.5B≤ 9B≤ 18B
T416 GB≤ 0.8B≤ 6.5B≤ 13B≤ 26B
RTX 409024 GB≤ 1.3B≤ 10.5B≤ 21B≤ 42B
L40S48 GB≤ 2.8B≤ 22.5B≤ 45B≤ 90B
A100 80GB / H10080 GB≤ 4.8B≤ 38.5B≤ 77B≤ 154B

Source: arithmetic — (VRAM − 3 GB budget) ÷ bytes/param (16 / 2 / 1 / 0.5), run 2026-10-09. These are arithmetic ceilings, not measured training runs: real activations at larger batch × sequence shrink every column, the full-fine-tune column most of all. Inference-side fit totals (weights + KV + runtime + 10% margin) come from the VRAM calculator.

How to read it: the full-fine-tune column is a hard upper bound (activations make it worse, never better); the adapter columns hold up in practice for batch-1-to-small runs and degrade gracefully — drop the batch, checkpoint, and you stay in the cell.

What to actually rent

  • Rate arithmetic belongs on the fine-tuning economics page — it tracks observed provider rows for LoRA/QLoRA candidates (A100, H100, L40S, H200) and its cost driver applies here unchanged: run-hours × hourly rate, because a 20% cheaper GPU rarely compensates for a 2× slower iteration loop.
  • For side-by-side rates while you decide: GPU comparison and cheapest GPUs.
  • The budget-shaped version of this decision — renting capability rather than reflex — appears as an 8B fine-tuning example in the automated GPU comparison article.

FAQ

Can I fine-tune an 8B model on a 24GB GPU? LoRA yes: FP16 base 16.1 GB + adapter + activations fits (inference-side reference total at 4K: 20.0 GB from the engine). Full fine-tuning no: 128.5 GB of training state — five times the card.

Is QLoRA worse than LoRA? Different frozen base precision, identical adapter training. You inherit INT4's 3–5% perplexity trade-off from the quantization explainer; whether that matters is a benchmark question about your data, not a rule.

Does gradient checkpointing make full fine-tuning fit on one GPU? No. Checkpointing shrinks the activation term; it does not touch the 16-bytes-per-parameter optimizer stack. An 8B full fine-tune is still 128.5 GB before activations.

What fits a rented 80GB card? Arithmetic ceilings: full fine-tune ≤ 4.8B, LoRA to 38.5B (FP16) or 77B (INT8), QLoRA to 154B — with activations eating into each. Verify inference-side base fits on the calculator.

Limitations and assumptions

  • Training figures are arithmetic, not measurements. Every total = parameter count × bytes/param with the formulas shown above; all are CALCULATED_ESTIMATES (confidence: arithmetic), run 2026-10-09. No training throughput, convergence, or loss numbers appear in this article — we have not measured them, and this site does not publish benchmark claims it can't source.
  • Conditions: AdamW mixed precision; 3 GB runtime/adapter/activation budget for ceilings (CUDA context + adapter + modest batch-1 activations); activations excluded from the Direct-answer table and scale linearly with batch × sequence.
  • Engine figures (20.0 GB for 8B FP16 at 4K, 40.8 GB for 70B INT4, 74.2 GB for 32B FP16) are inference-side totals from the VRAM canonical engine — weights + KV-cache + 1.2 GB runtime + 0.4 GB activation + 10% fragmentation margin, batch 1, single GPU — the same code path as the VRAM tiers article.
  • Not covered: multi-GPU sharding libraries, CPU/RAM offload, dataset-pipeline memory, LoRA-variant extras (QLoRA's paged optimizers, DoRA), and FP8 training. The single multi-GPU figure (16×80 GB for 70B full fine-tune) is aggregate arithmetic, not a deployment recipe.

Related resources