LLM Inference GPU Economics: Cost per Hour (2026)
Cheapest verified cloud GPUs for serving Llama-class models: observed hourly rates, VRAM fit, and the decode-bandwidth limit that decides tokens/sec.
The fast answer
Cheapest verified GPU: GeForce RTX 4090 at $0.34/hr on-demand (Vast.ai) among 6 candidate GPUs. Lowest on-demand hourly row for this GPU's providers.json key; refreshed daily.Observed On-Demand
VRAM needed: Llama 3.3 70B Instruct at INT4 needs 48.1 GB full-stack (weights 35 GB + KV-cache + overhead) at 128,000 tokens.
Cost driver: Memory bandwidth per dollar: the cheapest GPU that fits the model at the target context usually leads, but a bandwidth-starved card stretches latency at load.
Candidate GPUs for LLM Inference
Fit = full-stack VRAM total for Llama 3.3 70B Instruct (INT4/FP16 at model context) ≤ GPU VRAM. Rates are observed rows from data/providers.json, refreshed daily.
| GPU | VRAM | Bandwidth | INT4 fit | FP16 fit | On-demand | Spot |
|---|---|---|---|---|---|---|
| H200 SXM5 | 141 GB | 4.8 TB/s | ✓ fits | OOM | $2.79/hr | $2.79/hr |
| H100 SXM5 | 80 GB | 3.35 TB/s | ✓ fits | OOM | $1.89/hr | $1.89/hr |
| B200 Blackwell | 192 GB | 8.0 TB/s | ✓ fits | ✓ fits | $3.99/hr | $3.99/hr |
| A100 80GB SXM4 | 80 GB | 2.0 TB/s | ✓ fits | OOM | $1.59/hr | $1.59/hr |
| L40S | 48 GB | 864 GB/s | OOM | OOM | $0.69/hr | $0.69/hr |
| GeForce RTX 4090 | 24 GB | 1.0 TB/s | OOM | OOM | $0.34/hr | $0.34/hr |
Reference models for this workload
Methodology & provenance
Hourly rates are the lowest observed on-demand and spot rows for each GPU key in data/providers.json (refreshed daily by telemetry cron). VRAM fit uses the models-registry weight model (params × bytes-per-parameter) plus KV-cache sizing at the model's context window. Throughput ordering follows the decode-time bottleneck: token generation is memory-bandwidth bound, so ranking uses each GPU's spec-sheet bandwidth — no throughput figure is asserted without a benchmark row.
Next steps