Workload GPU Economics: Cost per Hour (2026)
Data-first cost guides for real workloads: observed provider rates, deterministic VRAM math from the models registry, and provenance labels on every figure. Pick a workload to see the cheapest verified GPU and which cards actually fit the model.
| Workload | Reference model | VRAM (INT4) | Cheapest verified GPU | On-demand |
|---|---|---|---|---|
| LLM Inference llm-inference | Llama 3.3 70B Instruct | 40 GB | GeForce RTX 4090 | $0.34/hr |
| RAG Pipeline rag-pipeline | Llama 3.1 8B Instruct | 6 GB | GeForce RTX 4090 | $0.34/hr |
| LLM Fine-Tuning llm-fine-tuning | Llama 3.3 70B Instruct | 40 GB | L40S | $0.69/hr |
| Image Generation image-generation | FLUX.1 [dev] | 12 GB | GeForce RTX 4090 | $0.34/hr |
| Batch Inference batch-inference | DeepSeek V3 | 336 GB | GeForce RTX 4090 | $0.34/hr |
Rates are observed on-demand rows from data/providers.json, refreshed daily (UTC). VRAM figures come from the canonical VRAM engine — weights + KV-cache + overhead. Every workload page documents its methodology.
Pre-Configured Cloud Clusters
Ready-to-deploy clusters optimized for specific AI workloads with verified configurations and pricing.
Optimized for DeepSeek R1 671B reasoning workloads
High-throughput embedding model serving
Optimized for FLUX.1 [dev] image generation
Speech-to-text transcription at scale
Cost-effective DeepSeek V2 236B serving
Production-ready Llama 3 70B serving with vLLM
Efficient Mistral Nemo 12B inference workloads
Image Generation Workload
Link to the image generation workload page.