Batch

Batch Inference GPU Economics: Offline Throughput Cost (2026)

Offline batch inference economics: why the cheapest large-VRAM GPU leads when latency is free and preemption risk is priced in.

Cheapest verified$0.34/hr
Reference modelDeepSeek V3
VRAM (INT4)336 GB
Candidates6 GPUs

The fast answer

Cheapest verified GPU: GeForce RTX 4090 at $0.34/hr on-demand (Vast.ai) among 6 candidate GPUs.

Observed On-Demand
Observed On-Demand RateHIGH
SourceObserved provider API rate (Vast.ai)
VerifiedSep 30, 2026
Value
0.34 USD/hr
Methodology

Lowest on-demand hourly row for this GPU's providers.json key; refreshed daily.

Refreshed daily from provider APIs and market scrapingSep 30, 2026

VRAM needed: DeepSeek V3 at INT4 needs 381.0 GB full-stack (weights 336 GB + KV-cache + overhead) at 128,000 tokens.

Cost driver: Tokens per GPU-hour at full VRAM utilization: batch workloads should pick the largest model that fits the card before comparing rates.

Candidate GPUs for Batch Inference

Fit = full-stack VRAM total for DeepSeek V3 (INT4/FP16 at model context) ≤ GPU VRAM. Rates are observed rows from data/providers.json, refreshed daily.

GPUVRAMBandwidthINT4 fitFP16 fitOn-demandSpot
B200 Blackwell192 GB8.0 TB/sOOMOOM$3.99/hr$3.99/hr
H100 SXM580 GB3.35 TB/sOOMOOM$1.89/hr$1.89/hr
A100 80GB SXM480 GB2.0 TB/sOOMOOM$1.59/hr$1.59/hr
H200 SXM5141 GB4.8 TB/sOOMOOM$2.79/hr$2.79/hr
L40S48 GB864 GB/sOOMOOM$0.69/hr$0.69/hr
GeForce RTX 409024 GB1.0 TB/sOOMOOM$0.34/hr$0.34/hr

Reference models for this workload

Methodology & provenance

Batch jobs trade latency for utilization: large batches push every GPU toward its bandwidth/compute ceiling, so cost-per-token ordering uses observed hourly rates against VRAM fit (registry weight model + KV-cache at batch context). Spot rows are shown where they exist because batch workloads can checkpoint and requeue — preemption tolerance is a workload property, not a pricing assumption.

Rates: observed provider API rows, refreshed daily (UTC).VRAM: canonical VRAM engine (weights + KV-cache + overhead + headroom).Full methodology →

Next steps

Frequently Asked Questions

What is the cheapest GPU for Batch Inference?▾
GeForce RTX 4090 at $0.34/hr on-demand (Vast.ai) is the lowest observed rate among the candidate GPUs for this workload. Rates refresh daily from provider APIs.
How much VRAM does Batch Inference need?▾
DeepSeek V3 as the reference model needs 336 GB at INT4 / 1340 GB at FP16 for weights; full-stack totals including KV-cache are 381.0 GB (INT4) and 1498.4 GB (FP16) at 128000 tokens context.
What drives cost for Batch Inference?▾
Tokens per GPU-hour at full VRAM utilization: batch workloads should pick the largest model that fits the card before comparing rates. All rates on this page are observed provider rows — never estimates.