⚡Under $0.50/hr🧠VRAM Estimator⚖Compare GPUs🎁Free LLM APIs🎯Model Index

Find the cheapest way to run your AI workload

Calculate VRAM for your model, compare verified spot GPU pricing across 10+ clouds, and find zero-cost API endpoints.

Lowest H100$1.89/hr
RTX 4090 Spot$0.34/hr
● Prices & Free Tiers Verified Daily (UTC)
🧠 What Are You Trying to Run?Instant Estimate
Weights VRAM
~185 GB
KV-Cache (8K)
~11.1 GB
Total VRAM (incl. 2GB CUDA)
~198 GB
Recommended GPU
2x H200 141GB
Lowest Verified Spot
$2.19/hr
View all H200 prices →
Free Serverless Alternative
DeepSeek API
$0.55/$2.19 per M tokens →
Deep Dive Profile
DeepSeek R1 (671B MoE)
Full benchmark →
ProviderGPU & VRAMInterconnectSpot RateOn-DemandMonthlyStatusAction
Community
PCIe 4.0 (64 GB/s)$0.34 / hr$0.85 / hr$208 / moInstant
Deploy →
Bare Metal
PCIe 4.0 (64 GB/s)$0.69 / hr$1.73 / hr$422 / moInstant
Deploy →
Community
PCIe 4.0 (64 GB/s)$0.69 / hr$1.73 / hr$422 / moInstant
Deploy →
Cloud
PCIe 4.0 (64 GB/s)$0.74 / hr$1.85 / hr$453 / moInstant
Deploy →
Dedicated
PCIe 4.0 (64 GB/s)$0.89 / hr$2.23 / hr$545 / moInstant
Deploy →
Cloud
PCIe 4.0 (64 GB/s)$1.09 / hr$2.73 / hr$667 / moInstant
Deploy →
Bare Metal
PCIe 4.0 (64 GB/s)$1.19 / hr$2.97 / hr$728 / moInstant
Deploy →
Dedicated
PCIe 4.0 (64 GB/s)$1.49 / hr$3.73 / hr$912 / moInstant
Deploy →
Dedicated
N/A$1.59 / hr$3.98 / hr$973 / moInstant
Deploy →
Community
NVLink 4.0 (900 GB/s)$1.89 / hr$4.72 / hr$1,157 / moInstant
Deploy →
Bare Metal
NVLink 4.0 (900 GB/s)$2.29 / hr$5.73 / hr$1,401 / moInstant
Deploy →
Community
NVLink 4.0 (900 GB/s)$2.79 / hr$6.98 / hr$1,707 / moInstant
Deploy →
Dedicated
NVLink 4.0 (900 GB/s)$2.99 / hr$7.48 / hr$1,830 / moInstant
Deploy →
Bare Metal
NVLink 4.0 (900 GB/s)$3.19 / hr$7.98 / hr$1,952 / moInstant
Deploy →
Cloud
NVLink 4.0 (900 GB/s)$3.49 / hr$8.73 / hr$2,136 / moInstant
Deploy →
Community
NVLink 5.0 (1.8 TB/s)$3.99 / hr$9.98 / hr$2,442 / moInstant
Deploy →
Dedicated
NVLink 4.0 (900 GB/s)$3.99 / hr$9.98 / hr$2,442 / moInstant
Deploy →
Cloud
NVLink 4.0 (900 GB/s)$4.31 / hr$10.77 / hr$2,638 / moInstant
Deploy →
Bare Metal
NVLink 5.0 (1.8 TB/s)$4.49 / hr$11.23 / hr$2,748 / moInstant
Deploy →
Dedicated
NVLink 5.0 (1.8 TB/s)$5.49 / hr$13.73 / hr$3,360 / moInstant
Deploy →
Cloud
NVLink 5.0 (1.8 TB/s)$5.99 / hr$14.98 / hr$3,666 / moInstant
Deploy →
Data Freshness: Public Cloud APIs & Market Scraping | Refreshed Daily (UTC)Benchmark Baseline: Ubuntu 24.04, CUDA 12.4, vLLM v0.6.x, PagedAttention v2, FlashAttention-3

Can I Run It? Model GPU Requirements

Check VRAM requirements, quantization limits, and recommended cloud GPUs for every supported model.

GPU Comparisons

Side-by-side technical comparisons: interconnect latency, virtualization overhead, storage I/O, and spot eviction risk.

VSH100

Spheron vs RunPod

Spheron offers dedicated bare-metal clusters with RDMA InfiniBand for multi-node distributed training.

Read Analysis →
VSRTX 4090

RunPod vs Vast.ai

Vast.ai offers the lowest spot pricing for constrained research and non-critical batch workloads.

Read Analysis →
VSH100

Spheron vs CoreWeave

CoreWeave provides massive-scale 1,000+ GPU HGX clusters for frontier foundation model training.

Read Analysis →
VSH100

Lambda Labs vs RunPod

Lambda Labs provides a pre-installed ML stack for PyTorch development.

Read Analysis →
VSH100

Hyperstack vs Verda

Verda is the standard deployment profile for EU data sovereignty requirements.

Read Analysis →
VSA100

Massed Compute vs RunPod

Massed Compute delivers predictable physical datacenter control.

Read Analysis →
VSH100

AWS EC2 P5 vs Specialist AI Clouds

AWS P5 is standard infrastructure if your pipeline relies on existing AWS IAM and S3.

Read Analysis →
VSH100

GCP A3 vs Lambda Labs

GCP A3 provides larger scale for Google ecosystem users, but Lambda Labs delivers lower overhead and better $/FLOP for independent teams..

Read Analysis →
VSH100

H100 SXM5 vs H200 SXM5

The H200 eliminates 2-GPU tensor sharding for Llama 70B at 128k context. The H100 remains the lower-cost option for sub-32k workloads at $1.89/hr vs $2.50/hr..

Read Analysis →
VSB200

H100 SXM5 vs B200 Blackwell

The B200 delivers 2x inference throughput via FP4 but requires liquid-cooled infrastructure at 1000W TDP. The H100 is deployment-ready in air-cooled datacenters today..

Read Analysis →
VSRTX 4090

RTX 4090 vs L40S

The RTX 4090 offers lower hourly cost for INT4 inference. The L40S leads on FP8 precision, memory capacity, enterprise reliability, and ECC memory for production workloads..

Read Analysis →

Hardware Radar & Price Trackers

Dedicated pages for every tracked GPU model with live pricing, architecture breakdowns, and benchmark data.

NEW

AI Compute 101 — GPU Fundamentals for LLM Deployment

Learn VRAM sizing, quantization formats, KV-cache scaling, and multi-GPU parallelism from first principles.

Start Learning →

Model Hosting & VRAM Sizing

Find the cheapest cloud GPU for your specific model. VRAM sizing, quantization overhead, and production runbooks.

Engineering Guides

Production architecture runbooks: network topologies, inference deployment, spot fault tolerance, and market telemetry.

LLM Token Pricing & Economics

Compare API token costs, self-hosted break-even thresholds, and routing gateway economics.

Latest Compute Intelligence & Guides

Curated how-to guides and research reports for AI infrastructure deployment.

Compute Research & News Reports