RAG Pipeline GPU Economics: Retrieval + Generation Cost (2026)
What a retrieval-augmented generation stack costs per hour: embedding retrieval on shared hardware, generation leg priced on verified GPU rates.
The fast answer
Cheapest verified GPU: GeForce RTX 4090 at $0.34/hr on-demand (Vast.ai) among 4 candidate GPUs. Lowest on-demand hourly row for this GPU's providers.json key; refreshed daily.Observed On-Demand
VRAM needed: Llama 3.1 8B Instruct at INT4 needs 14.8 GB full-stack (weights 4 GB + KV-cache + overhead) at 128,000 tokens.
Cost driver: Prompt tokens dominate RAG spend: longer retrieved contexts inflate prefill compute and KV-cache VRAM, which is why context length โ not just model size โ is the sizing knob.
Candidate GPUs for RAG Pipeline
Fit = full-stack VRAM total for Llama 3.1 8B Instruct (INT4/FP16 at model context) โค GPU VRAM. Rates are observed rows from data/providers.json, refreshed daily.
| GPU | VRAM | Bandwidth | INT4 fit | FP16 fit | On-demand | Spot |
|---|---|---|---|---|---|---|
| L40S | 48 GB | 864 GB/s | โ fits | โ fits | $0.69/hr | $0.69/hr |
| GeForce RTX 4090 | 24 GB | 1.0 TB/s | โ fits | OOM | $0.34/hr | $0.34/hr |
| A100 80GB SXM4 | 80 GB | 2.0 TB/s | โ fits | โ fits | $1.59/hr | $1.59/hr |
| H100 SXM5 | 80 GB | 3.35 TB/s | โ fits | โ fits | $1.89/hr | $1.89/hr |
Reference models for this workload
Methodology & provenance
RAG cost is dominated by the generation leg โ embedding and rerank models are weight-small compared with the answer model and run alongside it on the same GPU in practice. This page prices the generation leg from registry VRAM requirements (8B FP16 โ 16 GB, 32B INT4 โ 20 GB) against observed provider rates, and adds index-build context in the prompt budget via the calculator's context-length control.
Next steps