Cost Analysis2026-10-05โ€ขBy Sreeโ€ข5 min read

How Much Does It Cost to Run DeepSeek R1?

DeepSeek R1 costs across API, cloud GPU, and self-hosting. Why the 671B MoE model costs differently than it appears.

Direct answer

DeepSeek R1 costs fall into three distinct tiers depending on how you run it:

ApproachModel variantCost per 1M tokensNotes
DeepSeek APIDeepSeek R1 (671B MoE)$0.55 in / $2.19 outOfficial endpoint (30 tok/s API rate)
Self-hosting (H200)R1 INT4, 4ร— H200 cluster~$31$11.16/hr Vast.ai; 417 GB needed of 564 GB; 100 tok/s cluster estimate
Self-hosting (H200)R1 FP8, 6ร— H200 clusternot verifiable$16.74/hr; 786 GB needed; no published throughput record
API alternativeR1 Distill 70B$0.35โ€“$0.55 inDistilled variant โ€” most of the reasoning quality at a fraction of the size

The key insight: R1 is a Mixture-of-Experts model โ€” 671 billion total parameters, 37 billion active per token. The full weight matrix must be resident regardless of which experts fire: 671 GB at FP8, 336 GB at INT4, plus a 41.9 GB KV cache at 128K context (786 GB / 417 GB in service). That cluster cost, not the API price, dominates the economics.

The VRAM calculator computes exact memory requirements, and the DeepSeek model page lists verified API providers with current pricing.

Why DeepSeek R1 costs more than the API price suggests

The DeepSeek API charges $0.55/M input tokens. Self-hosting the same model means paying for a cluster 24/7 regardless of token volume. Take the 4ร— H200 INT4 config at $11.16/hr (Vast.ai):

Breakeven vs $0.55/M input: $11.16/hr รท ($0.55 per 1M) = 20.3M tokens/hour โ‰ˆ 5,640 tok/s sustained
Breakeven vs $0.88/M blended (4:1 in:out): 12.7M tokens/hour โ‰ˆ 3,523 tok/s sustained

Against an 8ร— H200 FP8 cluster at $22.32/hr the numbers double. Sustaining thousands of decode tokens per second means running with large batch concurrency across the whole cluster โ€” far beyond single-stream serving. For most workloads the API is cheaper per token; self-hosting is about data control, fine-tuning access, or removing per-request rate limits.

VRAM requirements make the decision for you

DeepSeek R1 VRAM (canonical engine, 128K context, batch=1):

PrecisionVRAM neededGPU requirement
FP161,524 GB12ร— H200 (1,692 GB) or two 8-GPU nodes
FP8786 GB6ร— H200 (846 GB) minimum ยท 10ร— H100 (800 GB, dual-node)
INT4417 GB3ร— H200 (423 GB, tight) ยท 4ร— H200 (564 GB, comfortable) ยท 6ร— H100 (480 GB)

At short context (โ‰ค4K) the FP8 total drops to 741 GB โ€” but even then 8ร— H100 (640 GB) does not fit FP8. The minimum viable cluster article covers cluster configuration in detail. The DeepSeek R1 model page shows the full GPU compatibility matrix.

API cost breakdown

From the model registry โ€” verified API providers:

ProviderInput /MOutput /MAPI tok/sNotes
DeepSeek (official)$0.55$2.1930Direct from DeepSeek
NVIDIA NIM$0.10$0.15500Requires NGC account
GitHub Models$0.15$0.2030Trial terms apply
OpenRouter$0.70$2.5022Aggregator
Together AI$0.80$2.4025Higher latency

Observation: NVIDIA NIM's $0.10/M input is the cheapest listed rate, with rate limits and an NGC account requirement. Always confirm current rates on the DeepSeek R1 model page.

Cloud GPU costs

Running DeepSeek R1 on cloud GPUs (INT4, all prices Vast.ai from providers.json):

Config$/hrVRAM availableVRAM needed (INT4 @128K)Throughput basis$/M tokens
4ร— H200 INT4$11.16564 GB417 GB โœ…100 tok/s (models.json cluster estimate)$31.00
8ร— H100 INT4$15.12640 GB417 GB โœ…no published recordnot verifiable
8ร— H200 INT4$22.321,128 GB417 GB โœ…no published recordnot verifiable
3ร— B200 INT4$11.97576 GB417 GB โœ…no published recordnot verifiable

Token throughput for multi-GPU R1 is [PROJECTED] territory โ€” MoE decode is memory-bandwidth bound, and no single public benchmark exists for 8ร— H200 R1 inference. The one throughput figure we do publish (100 tok/s) comes from the site's own cluster record in data/models.json (360,000 tokens/hr for the 4ร— H200 config). Treat every other cell as unverified rather than trusting FP8-TFLOPS linear scaling, which does not describe decode.

The distilled model alternative

DeepSeek publishes distilled variants that run on much smaller hardware:

ModelParams (B)VRAM (INT4)API tok/sAPI cost /M (input)
R1 Distill Llama 70B7035 GB35โ€“40$0.35โ€“$0.55
R1 Distill Qwen 32B3216 GB45โ€“50$0.25โ€“$0.35

The distilled variants retain much of the reasoning quality at a fraction of the hardware cost. The DeepSeek R1 model page has the complete API pricing comparison.

When self-hosting DeepSeek R1 makes sense

  1. Data privacy requirements prevent API usage
  2. Very high token volume โ€” tens of millions of tokens per hour sustained, with real batch concurrency (see the breakeven math above)
  3. Existing GPU infrastructure โ€” you already own or have reserved H200-class hardware running below full utilization
  4. Fine-tuning needs โ€” you must modify the weights

For most use cases, the DeepSeek API at $0.55/M input is the most cost-effective and operationally simple option. See the cheapest GPUs page for live H100/H200 rates if you're evaluating the self-hosting path.

Limitations

  • Multi-GPU R1 throughput is [PROJECTED] โ€” no single public benchmark exists for 8ร— H200 MoE inference
  • Cloud GPU pricing is [OBSERVED] as of 2026-10-05 and varies by provider
  • NVIDIA NIM pricing requires an NGC account and is subject to rate limits
  • VRAM figures already include the MLA-compressed KV cache (41.9 GB at 128K) โ€” they are not weights-only numbers
  • Power, storage, and facility costs are not included in self-hosting estimates

Related resources