Vast.ai — Peer-to-Peer GPU Marketplace
Workload Suitability Matrix: Multi-Node Training: Marketplace model aggregates idle consumer and enterprise GPUs; pricing determined by supply and demand; reliability varies by host with uptime rangin...
Cloud GPU Pricing
| GPU | VRAM | Interconnect | Spot | On-Demand | Monthly | Deploy |
|---|---|---|---|---|---|---|
| RTX 4090 | 24GB GDDR6X | PCIe 4.0 (64 GB/s) | $0.34 | $0.34 | $208 | Deploy → |
| L40S | 48GB GDDR6 | PCIe 4.0 (64 GB/s) | $0.69 | $0.69 | $422 | Deploy → |
| H100 SXM5 | 80GB HBM3 | NVLink 4.0 (900 GB/s) | $1.89 | $1.89 | $1,157 | Deploy → |
| H200 | 141GB HBM3e | NVLink 4.0 (900 GB/s) | $2.79 | $2.79 | $1,707 | Deploy → |
| B200 | 192GB HBM3e | NVLink 5.0 (1.8 TB/s) | $3.99 | $3.99 | $2,442 | Deploy → |
Competitor Comparison
| PROVIDER | MIN PRICE / HR | AVG TOKEN RATE | FREE TIER | BILLING | KEY ADVANTAGE |
|---|---|---|---|---|---|
| Vast.ai | $0.34 | — | No free tier; lowest hourly ra... | per-second | ★ You are here |
| RunPod | $0.39 | — | No free GPU tier; 20% off firs... | per-second | Community & Secure GPU Cloud |
Technical Nuances & Editorial Analysis
Workload Suitability Matrix: Multi-Node Training: Marketplace model aggregates idle consumer and enterprise GPUs; pricing determined by supply and demand; reliability varies by host with uptime ranging from partial to 99.9%. Spot Inference Prototyping: Lowest hourly rates across H100, H200, B200, and RTX 4090; community tier carries preemption risk; ideal for batch inference, fine-tuning, and experimentation. Persistent Production API: Users must verify GPU health and uptime history before committing to long-running workloads; CLI deployment and custom container images supported for vLLM, Ollama.
Billing Granularity
per-second
How charges are calculated
Hidden Costs
Variable reliability; some instances may be preempted by host
Watch out for these fees
Free Tier
No free tier; lowest hourly rates available
Credit requirements and caps
Best Use Case
Best for flexible GPU workloads
Ideal workload profile