Cloud GPU vs Self-Hosting: When Renting Actually Wins
Breakeven math for cloud GPU vs owning hardware: hourly cost, utilization, electricity, depreciation, and operational overhead. Worked example serving Llama 3.3 70B on H200.
Direct answer
Renting cloud GPU wins when utilization is below roughly 40% of the month, when you need elastic capacity without managing hardware, or when your workload is unpredictable. Owning wins when utilization sits consistently above ~45% and you have staff to run the infrastructure.
The crossover, with explicit inputs:
Monthly cloud cost (H100, Vast.ai): $1.89/hr ร hours used
Monthly self-hosting cost (example inputs): โ $575 fixed
depreciation $389 ($14,000 hardware รท 36 months โ input, not our data)
electricity $86 (1 kW system incl. host, $0.12/kWh, 24/7)
colo/rack $100
Breakeven: $575 รท $1.89/hr โ 304 hr/month โ 42% of the month (~10 hr/day)
If your workload would run the GPU more than ~42% of the month, owning is cheaper; below that, renting is. Notice the direction: idle time favors the cloud, busy time favors ownership. The exact crossover moves with purchase price, electricity rate, and whether you already pay for the rack.
The host vs API breakeven guide covers the API-vs-self-host variant of the same math.
Breaking down the true cost
Cloud GPU hourly cost
Spot/on-demand pricing fluctuates โ check live rates. Observed in providers.json (2026-10-03):
| GPU | Vast.ai | RunPod on-demand (spot) |
|---|---|---|
| H100 SXM5 | $1.89 | $3.49 ($2.49 spot) |
| H200 SXM5 | $2.79 | $4.31 |
| L40S | $0.69 | $1.09 |
| RTX 4090 | $0.34 | $0.74 ($0.39 spot) |
These rates are observed snapshots from providers.json, not live telemetry โ see the individual GPU pages for current values.
Self-hosting cost components
| Component | Monthly estimate | Status |
|---|---|---|
| Electricity | $86โ$120 | assumption โ 1 kW system (700W GPU + host) at $0.10โ$0.15/kWh, 24/7 |
| Hardware depreciation | input | your input โ our data tracks cloud rates only; divide your purchase quote by useful life |
| Rack/space fees | $50โ$150 | assumption โ colocation or home lab |
| Bandwidth/storage | $20โ$50 | assumption โ varies by connection |
| Worked-example total | $575 | = $389 depreciation ($14kรท36) + $86 electricity + $100 colo |
Utilization matters more than you think
A GPU you own costs money whether it's idle or busy. A cloud GPU costs only while it runs. Using the $575 fixed-cost worked example:
At 33% utilization (8 hr/day):
- Cloud H100: $1.89 ร 8 ร 30 = $454/month โ renting wins
- Self-hosted: $575/month fixed
At 79% utilization (19 hr/day):
- Cloud H100: $1.89 ร 19 ร 30 = $1,077/month
- Self-hosted: $575/month โ owning wins
The crossover is at ~42% utilization (โ10 hr/day at H100 $1.89/hr). Below it, rent. Above it, buy โ before counting operational overhead, which usually widens the owning case further only if your team is small and efficient.
Worked example: Llama 3.3 70B inference
Scenario: serve Llama 3.3 70B at FP8, 32K context, processing 2 million tokens per day.
Config note (this is the part older versions got wrong): at 32K, 70B FP8 needs 81 GB of service VRAM (71 GB weights + ~1 GB KV with the 28-layer preset + overhead). That does not fit a single H100 (80 GB). The correct single-GPU cloud config is an H200 (141 GB, $2.79/hr, ~135 tok/s site record) โ or 2ร H100 with TP=2 at $3.78/hr.
Cloud approach (1ร H200):
- Fits 70B FP8 at 32K with 60 GB headroom
- Throughput: ~135 tok/s (site record, batch=1)
- Active hours needed: 2M รท 135 tok/s = ~4.1 hours/day
- Cloud cost: $2.79 ร 4.1 ร 30 = $344/month
Self-hosting approach (worked inputs):
- Hardware purchase: input (our data does not track purchase prices; example uses $14,000 for an H200-class single-GPU system โ replace with your quote)
- Depreciation (36 months): $14,000 รท 36 = $389/month
- Electricity (1 kW system, 24/7, $0.12/kWh): $86/month
- Rack/colo fee: $100/month
- Total: $575/month fixed
At 2M tokens/day, cloud is cheaper ($344 vs $575). If you grow to 6M tokens/day:
- Cloud: $2.79 ร 12.4 hr ร 30 = $1,033/month
- Self-hosting: still $575/month
Breakeven โ 3.3M tokens/day โ above that, the owned GPU wins (before ops time).
Operational overhead: the hidden cost
Cloud GPU pricing does not include:
- System administration time โ patching, monitoring, scaling, debugging
- Uptime management โ dealing with spot instance preemptions
- Networking โ load balancing, TLS termination, rate limiting
- Storage โ model checkpoint storage, log aggregation
- Security โ access control, vulnerability management
If engineering time is valued at $100/hour and GPU infrastructure absorbs 5 hours/month, that's $500/month of hidden cost โ on either side of the ledger, but cloud providers absorb some of it into the hourly rate while ownership puts it on your payroll.
Self-hosting has the opposite cost structure: high upfront capital, low per-month variable, but it requires dedicated ops expertise.
When renting wins
- Unpredictable traffic โ spiky workloads don't justify fixed hardware
- Short project timelines โ 2โ6 month projects rarely amortize a purchase
- No ops team โ managed infrastructure is the product you're buying
- Experimentation phase โ testing many models at small scale
- Burst capacity โ baseline on owned hardware, burst to cloud during peaks
The calculator models VRAM and hourly costs, and includes a break-even table against API pricing (labeled CALCULATED_ESTIMATE there) โ pair it with the utilization math above.
When self-hosting wins
- Consistent utilization โ above ~42% of the month (10+ hr/day) with the worked inputs
- Data privacy constraints โ data cannot leave your perimeter
- Custom hardware requirements โ specific interconnect or memory configs
- Multiple concurrent workloads โ pack several models on the same hardware
- Long time horizon โ 2+ year deployment without forced migration
The cheapest GPUs page tracks spot and on-demand pricing across 4 providers (Spheron, RunPod, Vast.ai, Lambda) โ useful when comparing a mixed reservation strategy.
Limitations and assumptions
- Electricity assumes $0.12/kWh (US-average ballpark); regional rates vary widely
- Hardware purchase price is a user input โ this site's data covers cloud rental rates only
- Depreciation assumes a 36-month useful life; actual life varies
- Cloud spot pricing is volatile โ rates spike during peak demand
- Network egress charges are not included (can be significant for API-heavy workloads)
- Maintenance time is an assumption โ actual ops burden varies by team
- Throughput figures are site batch=1 records; production batching improves every cloud-side number
Related resources
- Host vs API Breakeven โ the API-vs-self-host variant
- VRAM calculator โ model VRAM and cost modeling
- Cheapest GPUs โ live spot pricing
- H100 SXM5 GPU Specs โ rental rates
Related Articles
How Much Does It Cost to Run DeepSeek R1?
DeepSeek R1 costs across API, cloud GPU, and self-hosting. Why the 671B MoE model costs differently than it appears.
Cost AnalysisFLUX.1 [dev] Cost-Per-Image: RTX 4090 vs L40S Pricing Breakdown
FLUX.1 [dev] VRAM sizing, quantization, and cloud spot rates for RTX 4090 vs L40S. Dollar-per-image math with explicit throughput assumptions, not hidden ones.