LLM Fine-Tuning GPU Economics: LoRA & QLoRA Cost (2026)
Verified hourly rates for LoRA and QLoRA fine-tuning: which GPUs fit 32B-70B adapters, and where optimizer state beats weights.
The fast answer
Cheapest verified GPU: L40S at $0.69/hr on-demand (Vast.ai) among 4 candidate GPUs. Lowest on-demand hourly row for this GPU's providers.json key; refreshed daily.Observed On-Demand
VRAM needed: Llama 3.3 70B Instruct at INT4 needs 48.1 GB full-stack (weights 35 GB + KV-cache + overhead) at 128,000 tokens.
Cost driver: Run-hours × hourly rate: fine-tuning cost scales with epochs and dataset size, so a 20% cheaper GPU rarely compensates for a 2x slower iteration loop.
Candidate GPUs for LLM Fine-Tuning
Fit = full-stack VRAM total for Llama 3.3 70B Instruct (INT4/FP16 at model context) ≤ GPU VRAM. Rates are observed rows from data/providers.json, refreshed daily.
| GPU | VRAM | Bandwidth | INT4 fit | FP16 fit | On-demand | Spot |
|---|---|---|---|---|---|---|
| A100 80GB SXM4 | 80 GB | 2.0 TB/s | ✓ fits | OOM | $1.59/hr | $1.59/hr |
| H100 SXM5 | 80 GB | 3.35 TB/s | ✓ fits | OOM | $1.89/hr | $1.89/hr |
| L40S | 48 GB | 864 GB/s | OOM | OOM | $0.69/hr | $0.69/hr |
| H200 SXM5 | 141 GB | 4.8 TB/s | ✓ fits | OOM | $2.79/hr | $2.79/hr |
Reference models for this workload
Methodology & provenance
Fine-tuning VRAM = quantized base weights (registry INT4 figures) + adapter/optimizer state + activations. QLoRA keeps the base frozen at INT4 (70B ≈ 38 GB from the registry weight model), so the fit test is 48 GB-class cards and up; full FP16 fine-tuning of a 70B (≈140 GB weights) is excluded from single-GPU candidates by design. Rates are observed provider rows; run-hours are the raw material of training cost — no cloud credit math is assumed.
Next steps