Fine-Tuning

LLM Fine-Tuning GPU Economics: LoRA & QLoRA Cost (2026)

Verified hourly rates for LoRA and QLoRA fine-tuning: which GPUs fit 32B-70B adapters, and where optimizer state beats weights.

Cheapest verified$0.69/hr
Reference modelLlama 3.3 70B Instruct
VRAM (INT4)40 GB
Candidates4 GPUs

The fast answer

Cheapest verified GPU: L40S at $0.69/hr on-demand (Vast.ai) among 4 candidate GPUs.

Observed On-Demand
Observed On-Demand RateHIGH
SourceObserved provider API rate (Vast.ai)
VerifiedSep 30, 2026
Value
0.69 USD/hr
Methodology

Lowest on-demand hourly row for this GPU's providers.json key; refreshed daily.

Refreshed daily from provider APIs and market scrapingSep 30, 2026

VRAM needed: Llama 3.3 70B Instruct at INT4 needs 48.1 GB full-stack (weights 35 GB + KV-cache + overhead) at 128,000 tokens.

Cost driver: Run-hours × hourly rate: fine-tuning cost scales with epochs and dataset size, so a 20% cheaper GPU rarely compensates for a 2x slower iteration loop.

Candidate GPUs for LLM Fine-Tuning

Fit = full-stack VRAM total for Llama 3.3 70B Instruct (INT4/FP16 at model context) ≤ GPU VRAM. Rates are observed rows from data/providers.json, refreshed daily.

GPUVRAMBandwidthINT4 fitFP16 fitOn-demandSpot
A100 80GB SXM480 GB2.0 TB/s✓ fitsOOM$1.59/hr$1.59/hr
H100 SXM580 GB3.35 TB/s✓ fitsOOM$1.89/hr$1.89/hr
L40S48 GB864 GB/sOOMOOM$0.69/hr$0.69/hr
H200 SXM5141 GB4.8 TB/s✓ fitsOOM$2.79/hr$2.79/hr

Reference models for this workload

Methodology & provenance

Fine-tuning VRAM = quantized base weights (registry INT4 figures) + adapter/optimizer state + activations. QLoRA keeps the base frozen at INT4 (70B ≈ 38 GB from the registry weight model), so the fit test is 48 GB-class cards and up; full FP16 fine-tuning of a 70B (≈140 GB weights) is excluded from single-GPU candidates by design. Rates are observed provider rows; run-hours are the raw material of training cost — no cloud credit math is assumed.

Rates: observed provider API rows, refreshed daily (UTC).VRAM: canonical VRAM engine (weights + KV-cache + overhead + headroom).Full methodology →

Next steps

Frequently Asked Questions

What is the cheapest GPU for LLM Fine-Tuning?▾
L40S at $0.69/hr on-demand (Vast.ai) is the lowest observed rate among the candidate GPUs for this workload. Rates refresh daily from provider APIs.
How much VRAM does LLM Fine-Tuning need?▾
Llama 3.3 70B Instruct as the reference model needs 40 GB at INT4 / 140 GB at FP16 for weights; full-stack totals including KV-cache are 48.1 GB (INT4) and 172.1 GB (FP16) at 128000 tokens context.
What drives cost for LLM Fine-Tuning?▾
Run-hours × hourly rate: fine-tuning cost scales with epochs and dataset size, so a 20% cheaper GPU rarely compensates for a 2x slower iteration loop. All rates on this page are observed provider rows — never estimates.