How Much Does It Cost to Run DeepSeek R1?
DeepSeek R1 costs across API, cloud GPU, and self-hosting. Why the 671B MoE model costs differently than it appears.
Direct answer
DeepSeek R1 costs fall into three distinct tiers depending on how you run it:
| Approach | Model variant | Cost per 1M tokens | Notes |
|---|---|---|---|
| DeepSeek API | DeepSeek R1 (671B MoE) | $0.55 in / $2.19 out | Official endpoint (30 tok/s API rate) |
| Self-hosting (H200) | R1 INT4, 4ร H200 cluster | ~$31 | $11.16/hr Vast.ai; 417 GB needed of 564 GB; 100 tok/s cluster estimate |
| Self-hosting (H200) | R1 FP8, 6ร H200 cluster | not verifiable | $16.74/hr; 786 GB needed; no published throughput record |
| API alternative | R1 Distill 70B | $0.35โ$0.55 in | Distilled variant โ most of the reasoning quality at a fraction of the size |
The key insight: R1 is a Mixture-of-Experts model โ 671 billion total parameters, 37 billion active per token. The full weight matrix must be resident regardless of which experts fire: 671 GB at FP8, 336 GB at INT4, plus a 41.9 GB KV cache at 128K context (786 GB / 417 GB in service). That cluster cost, not the API price, dominates the economics.
The VRAM calculator computes exact memory requirements, and the DeepSeek model page lists verified API providers with current pricing.
Why DeepSeek R1 costs more than the API price suggests
The DeepSeek API charges $0.55/M input tokens. Self-hosting the same model means paying for a cluster 24/7 regardless of token volume. Take the 4ร H200 INT4 config at $11.16/hr (Vast.ai):
Breakeven vs $0.55/M input: $11.16/hr รท ($0.55 per 1M) = 20.3M tokens/hour โ 5,640 tok/s sustained
Breakeven vs $0.88/M blended (4:1 in:out): 12.7M tokens/hour โ 3,523 tok/s sustained
Against an 8ร H200 FP8 cluster at $22.32/hr the numbers double. Sustaining thousands of decode tokens per second means running with large batch concurrency across the whole cluster โ far beyond single-stream serving. For most workloads the API is cheaper per token; self-hosting is about data control, fine-tuning access, or removing per-request rate limits.
VRAM requirements make the decision for you
DeepSeek R1 VRAM (canonical engine, 128K context, batch=1):
| Precision | VRAM needed | GPU requirement |
|---|---|---|
| FP16 | 1,524 GB | 12ร H200 (1,692 GB) or two 8-GPU nodes |
| FP8 | 786 GB | 6ร H200 (846 GB) minimum ยท 10ร H100 (800 GB, dual-node) |
| INT4 | 417 GB | 3ร H200 (423 GB, tight) ยท 4ร H200 (564 GB, comfortable) ยท 6ร H100 (480 GB) |
At short context (โค4K) the FP8 total drops to 741 GB โ but even then 8ร H100 (640 GB) does not fit FP8. The minimum viable cluster article covers cluster configuration in detail. The DeepSeek R1 model page shows the full GPU compatibility matrix.
API cost breakdown
From the model registry โ verified API providers:
| Provider | Input /M | Output /M | API tok/s | Notes |
|---|---|---|---|---|
| DeepSeek (official) | $0.55 | $2.19 | 30 | Direct from DeepSeek |
| NVIDIA NIM | $0.10 | $0.15 | 500 | Requires NGC account |
| GitHub Models | $0.15 | $0.20 | 30 | Trial terms apply |
| OpenRouter | $0.70 | $2.50 | 22 | Aggregator |
| Together AI | $0.80 | $2.40 | 25 | Higher latency |
Observation: NVIDIA NIM's $0.10/M input is the cheapest listed rate, with rate limits and an NGC account requirement. Always confirm current rates on the DeepSeek R1 model page.
Cloud GPU costs
Running DeepSeek R1 on cloud GPUs (INT4, all prices Vast.ai from providers.json):
| Config | $/hr | VRAM available | VRAM needed (INT4 @128K) | Throughput basis | $/M tokens |
|---|---|---|---|---|---|
| 4ร H200 INT4 | $11.16 | 564 GB | 417 GB โ | 100 tok/s (models.json cluster estimate) | $31.00 |
| 8ร H100 INT4 | $15.12 | 640 GB | 417 GB โ | no published record | not verifiable |
| 8ร H200 INT4 | $22.32 | 1,128 GB | 417 GB โ | no published record | not verifiable |
| 3ร B200 INT4 | $11.97 | 576 GB | 417 GB โ | no published record | not verifiable |
Token throughput for multi-GPU R1 is [PROJECTED] territory โ MoE decode is memory-bandwidth bound, and no single public benchmark exists for 8ร H200 R1 inference. The one throughput figure we do publish (100 tok/s) comes from the site's own cluster record in
data/models.json(360,000 tokens/hr for the 4ร H200 config). Treat every other cell as unverified rather than trusting FP8-TFLOPS linear scaling, which does not describe decode.
The distilled model alternative
DeepSeek publishes distilled variants that run on much smaller hardware:
| Model | Params (B) | VRAM (INT4) | API tok/s | API cost /M (input) |
|---|---|---|---|---|
| R1 Distill Llama 70B | 70 | 35 GB | 35โ40 | $0.35โ$0.55 |
| R1 Distill Qwen 32B | 32 | 16 GB | 45โ50 | $0.25โ$0.35 |
The distilled variants retain much of the reasoning quality at a fraction of the hardware cost. The DeepSeek R1 model page has the complete API pricing comparison.
When self-hosting DeepSeek R1 makes sense
- Data privacy requirements prevent API usage
- Very high token volume โ tens of millions of tokens per hour sustained, with real batch concurrency (see the breakeven math above)
- Existing GPU infrastructure โ you already own or have reserved H200-class hardware running below full utilization
- Fine-tuning needs โ you must modify the weights
For most use cases, the DeepSeek API at $0.55/M input is the most cost-effective and operationally simple option. See the cheapest GPUs page for live H100/H200 rates if you're evaluating the self-hosting path.
Limitations
- Multi-GPU R1 throughput is [PROJECTED] โ no single public benchmark exists for 8ร H200 MoE inference
- Cloud GPU pricing is [OBSERVED] as of 2026-10-05 and varies by provider
- NVIDIA NIM pricing requires an NGC account and is subject to rate limits
- VRAM figures already include the MLA-compressed KV cache (41.9 GB at 128K) โ they are not weights-only numbers
- Power, storage, and facility costs are not included in self-hosting estimates
Related resources
- DeepSeek R1 Model Page โ API pricing and VRAM requirements
- VRAM Calculator โ exact memory calculations for MoE models
- Cheapest GPUs โ live H100/H200 spot pricing
- Minimum Viable Cluster โ hardware sizing guide
Related Articles
Cloud GPU vs Self-Hosting: When Renting Actually Wins
Breakeven math for cloud GPU vs owning hardware: hourly cost, utilization, electricity, depreciation, and operational overhead. Worked example serving Llama 3.3 70B on H200.
Cost AnalysisFLUX.1 [dev] Cost-Per-Image: RTX 4090 vs L40S Pricing Breakdown
FLUX.1 [dev] VRAM sizing, quantization, and cloud spot rates for RTX 4090 vs L40S. Dollar-per-image math with explicit throughput assumptions, not hidden ones.