DeepSeek R1 Serving Guide: FP8 vs INT4 VRAM, vLLM Config, and H200/B200 Cost Analysis
DeepSeek R1 671B MoE serving guide: INT4 = 403 GB VRAM (4x H200, TP=4); FP8 = 799 GB (6x H200). vLLM configuration with bitsandbytes-nf4 quantization and breakeven analysis at $0.55/M API vs $11.16/hr self-hosted.
Direct answer
DeepSeek R1 (671B MoE, 37B active) fits on 4ร H200 with INT4 quantized weights (โ375 GB VRAM at 128K context with 10% fragmentation headroom, TP=4 recommended for even MoE expert split; 3ร H200 is the hard minimum at 423 GB) or 6ร H200 with FP8 (โ745 GB, TP=6). Using a more conservative 18% allocator headroom (the /calculator setting) yields โ403 GB INT4 / โ799 GB FP8 โ same GPU counts. For most workloads, an R1-compatible API is cheaper than self-hosting; self-hosting only becomes cost-competitive above ~305M tokens/day under a 24/7 rental assumption. Use vLLM 0.6+ with --tensor-parallel-size 4 (INT4) or --tensor-parallel-size 6 (FP8) and --max-model-len 163840 for full 160K context.
Why DeepSeek R1 deployment is different
DeepSeek R1 is a Mixture-of-Experts (MoE) model with 671B total parameters but only 37B active at inference time (5.5% of total). This means:
| Metric | Value | Source |
|---|---|---|
| Total parameters | 671B | models-registry.json:deepseek-r1 |
| Active parameters (per token) | 37B | models-registry.json:deepseek-r1 |
| Architecture | MoE + MLA | config.json |
| Context window | 160,000 tokens | config.json max_position_embeddings=163840 |
| Layers | 61 | config.json num_hidden_layers=61 |
| Attention heads | 128 | config.json num_attention_heads=128 |
| KV heads | 128 | config.json num_key_value_heads=128 |
| Head dimension | 128 | config.json |
| Routed experts | 64 | config.json n_routed_experts=64 |
| kvLoraRank | 512 | config.json kv_lora_rank=512 |
| qkRopeHeadDim | 64 | config.json qk_rope_head_dim=64 |
Unlike dense models, the 671B parameter count can be misleading โ only 37B are active per forward pass. However, all 671B weights must be loaded into memory (either VRAM or CPU/disk with offloading) because the model dynamically routes tokens to different experts.
DeepSeek R1 uses Multi-Latent Attention (MLA) rather than traditional Grouped-Query Attention (GQA). This compresses the KV cache significantly: instead of storing per-head keys and values (128 heads ร 128 dims ร 2 = 32,768 values/token/layer), MLA stores a compressed latent of kvLoraRank (512) + qkRopeHeadDim (64) = 576 dimensions per token per layer โ a 57ร reduction in KV cache memory.
Benchmarks (from models-registry.json):
| Benchmark | DeepSeek R1 | DeepSeek R1 Distill 70B | GPT-4o |
|---|---|---|---|
| SweBench (coding) | 49.2 | 50.8 | 49.0 |
| LiveCodeBench | 65.9 | 60.1 | 42.0 |
| AIME 2024 (math) | 79.8 | 84.6 | 75.5 |
| Math500 | 97.3 | 92.5 | 75.5 |
| MMLU (general knowledge) | 90.8 | 88.0 | 86.4 |
R1 leads on math and coding benchmarks among open-weight models, which justifies the deployment complexity for reasoning-heavy workloads.
VRAM requirements by quantization level
All VRAM numbers include KV cache for 128K context, MLA KV cache for 160K is 5.36 GB (FP8/INT4/INT8) or 10.72 GB (FP16). Sources: models-registry.json, config.json, and deepseek-r1-fp8-vs-int4-memory. The calculator engine at /calculator uses 18% fragmentation as a conservative estimate; the canonical engine uses 10% (matching vLLM's allocator behavior).
Memory breakdown components
| Component | Description | Value (DeepSeek R1) |
|---|---|---|
| Model weights | Static parameters loaded in VRAM. FP16 = 2 bytes/param, FP8 = 1 byte/param, INT4 = 0.5 bytes/param. MoE means 671B total but only 37B active per token. | FP8: 671 GB, INT4: 336 GB, FP16: 1,342 GB |
| KV cache | Attention context storage per request. MLA compresses to kvLoraRank (512) + qkRopeHeadDim (64) = 576 dims per token per layer. Stored at 1 byte/element for FP8/INT4/INT8, 2 bytes for FP16. | FP8/INT4: 4.19 GB @ 128K, FP16: 8.38 GB @ 128K |
| CUDA runtime overhead | Fixed PyTorch/vLLM CUDA context, memory allocator buffers. | 1.2 GB (fixed) |
| Activation memory | Autoregressive decode workspace for batch=1. | 0.4 GB (fixed) |
| Fragmentation headroom | 10โ18% safety margin for memory allocator fragmentation. | 10% (canonical) or 18% (calculator) |
VRAM totals at 128K context
| Quantization | Weights | KV Cache (MLA) | CUDA | Activations | Fragmentation | Total VRAM |
|---|---|---|---|---|---|---|
| FP16 | 1,342 GB | 8.38 GB | 1.2 GB | 0.4 GB | 136.0 GB (10%) | 1,487.2 GB |
| FP8 | 671 GB | 4.19 GB | 1.2 GB | 0.4 GB | 68.0 GB (10%) | 744.5 GB |
| INT4 | 335.5 GB | 4.19 GB | 1.2 GB | 0.4 GB | 34.3 GB (10%) | 375.4 GB |
| INT8 | 671 GB | 4.19 GB | 1.2 GB | 0.4 GB | 68.0 GB (10%) | 744.5 GB |
Note: FP8 and INT8 produce identical VRAM totals (both use 1 byte/param for weights). INT4 is the only quantization that meaningfully reduces VRAM. KV cache is precision-invariant for FP8/INT4/INT8 (1 byte/element) and doubles for FP16 (2 bytes/element).
How many GPUs do you need? (128K context, 18% fragmentation)
| GPU | VRAM (each) | FP8 (799 GB) | INT4 (403 GB) | INT8 (799 GB) |
|---|---|---|---|---|
| H200 SXM5 (141 GB) | 141 GB | 6 GPUs | 3 GPUs | 6 GPUs |
| B200 (192 GB) | 192 GB | 5 GPUs | 3 GPUs | 5 GPUs |
| H100 SXM5 (80 GB) | 80 GB | 10 GPUs | 6 GPUs | 10 GPUs |
| RTX 4090 (24 GB) | 24 GB | โ | โ | โ |
Recommendation: Use 4ร H200 with TP=4 for INT4 serving. This provides even MoE expert distribution (64 experts รท 4 ranks = 16 per rank) and 161 GB headroom over the 403 GB requirement. For FP8, use 6ร H200 with TP=6 โ sufficient VRAM (846 GB vs 799 GB needed). For even MoE distribution with FP8, TP=8 provides 64 รท 8 = 8 experts per rank, at the cost of 2 additional GPUs.
Cost analysis: API vs self-hosted
DeepSeek API pricing
| Provider | Input | Output | Throughput | Source |
|---|---|---|---|---|
| DeepSeek | $0.55/M | $2.19/M | 30 tok/s | models-registry.json |
| NVIDIA NIM | $0.10/M | $0.15/M | 500 tok/s | models-registry.json |
| GitHub Models | $0.15/M | $0.20/M | 30 tok/s | models-registry.json |
| OpenRouter | $0.70/M | $2.50/M | 22 tok/s | models-registry.json |
| Together AI | $0.80/M | $2.40/M | 25 tok/s | models-registry.json |
| Fireworks | $0.90/M | $2.50/M | 28 tok/s | models-registry.json |
Note: As of 2026-10-11, DeepSeek R1 is no longer listed on DeepSeek's pricing page. The $0.55/M rate reflects historical DeepSeek API pricing; Together, OpenRouter, Fireworks, GitHub Models, and NIM rates are registry-sourced third-party prices. Verify current pricing with each provider before deployment โ do not treat this table as a live price feed.
Self-hosting costs (cloud GPU rates)
| GPU | Provider | Hourly | Monthly (720h) |
|---|---|---|---|
| H200 SXM | Vast.ai | $2.79 | $1,707.20 |
| H200 SXM | Spheron | $3.19 | $1,952.00 |
| H200 SXM | RunPod | $4.31 | $2,638.40 |
| B200 SXM | Vast.ai | $3.99 | $2,441.60 |
| B200 SXM | Spheron | $4.49 | $2,748.80 |
| H100 SXM | Vast.ai | $1.89 | $1,076.80 |
All provider rates from H200 cloud pricing, verified 2026-10-10.
Breakeven calculation
Assumptions:
- 4ร H200 INT4 (TP=4) on Vast.ai spot: 4 ร $2.79 = $11.16/hour
- 24-hour continuous operation: $11.16 ร 24 = $267.84 per day
- Local throughput: 120 tokens/sec (realistic for TP=4 with INT4, vLLM 0.6+)
- Daily throughput: 120 ร 3,600 ร 24 = 10.37M tokens/day
- API token mix: 80% input, 20% output (standard 4:1 ratio for RAG/codegen)
API average cost (DeepSeek, 4:1 ratio):
- $0.55 ร 0.8 + $2.19 ร 0.2 = $0.878 per million tokens
Self-hosted cost per million tokens (4ร H200, INT4, 120 tok/s):
- $11.16 / (120 ร 3,600 / 1,000,000) = $25.84 per million tokens
Breakeven (24/7 rental assumption): self-hosted daily cost is a fixed $267.84 whether you use 1M tokens or 500M โ you pay for idle GPUs. API cost scales linearly with volume at $0.878/M. Self-hosting becomes cheaper when:
daily_tokens ร $0.878 / 1,000,000 > $267.84
daily_tokens > ~305M
| Daily token volume | API cost (DeepSeek, 4:1 mix) | Self-hosted 4ร H200 (24/7) | Cheaper option |
|---|---|---|---|
| 1M tokens | $0.88 | $267.84 | API |
| 10M tokens | $8.78 | $267.84 | API |
| 50M tokens | $43.90 | $267.84 | API |
| 100M tokens | $87.80 | $267.84 | API |
| 200M tokens | $175.60 | $267.84 | API |
| 305M tokens | $267.79 | $267.84 | breakeven |
| 500M tokens | $439.00 | $267.84 | Self-hosted |
If you rent GPUs only for the hours you actually generate (no idle time), self-hosted cost is $25.84 per million tokens at 120 tok/s โ still far above the API's $0.878/M, so API wins at every volume. The 305M/day breakeven exists only under 24/7 committed rental (or reserved instances) where idle capacity is sunk cost. Pick the framing that matches how you'd actually procure GPUs.
Breakeven summary: Self-hosting on 4ร H200 (INT4, TP=4) at 120 tok/s becomes cheaper than the DeepSeek API above ~305M tokens/day โ but only if the cluster runs 24/7. Under on-demand, pay-for-what-you-use renting, the API is always cheaper per token. Self-hosting also wins earlier when you need privacy, deterministic latency, or freedom from API rate limits (see below).
| Throughput | Tokens/hr | Daily tokens | Daily self-hosted cost | Cost per M tokens (self-hosted) | Breakeven daily volume |
|---|---|---|---|---|---|
| 80 tok/s | 288K | 6.91M | $267.84 | $38.76 | 305M |
| 100 tok/s | 360K | 8.64M | $267.84 | $31.00 | 305M |
| 120 tok/s | 432K | 10.37M | $267.84 | $25.84 | 305M |
| 150 tok/s | 540K | 12.96M | $267.84 | $20.67 | 305M |
| 200 tok/s | 720K | 17.28M | $267.84 | $15.50 | 305M |
Note: The breakeven daily volume is the same regardless of throughput โ it's driven by the fixed daily rental cost ($267.84) and the API cost per million tokens ($0.878/M). Throughput only affects the per-token cost of self-hosting.
Key insight: DeepSeek's API is extremely cost-competitive ($0.878/M) compared to most providers. Only NVIDIA NIM ($0.11/M) and GitHub Models ($0.15/M) are cheaper per-token. Self-hosting only wins at very high daily volumes (>305M tokens) or when you need:
- Deterministic latency (no network round-trip)
- Data privacy (can't send prompts to external APIs)
- Unlimited rate limits (API caps at 30 tokens/sec โ max ~2.6M tokens/day per key)
- Custom model modifications (e.g., custom loss functions, LoRA adapters)
For comparison with other models:
| Model | API input | API output | Self-hosted est. | Notes |
|---|---|---|---|---|
| DeepSeek R1 | $0.55/M | $2.19/M | $15.50โ38.76/M (4x H200, 80โ200 tok/s) | API wins unless >305M tokens/day |
| Llama 3.3 70B | $0.88/M | $0.88/M | ~$12/M (1x H200) | Simpler deployment, lower VRAM |
| Qwen 2.5 72B | $0.20/M | $0.60/M | ~$12/M (1x H200) | Better cost-to-quality ratio |
See DeepSeek R1 model page for full API provider comparisons and cloud GPU pricing for live rate data.
vLLM configuration
Based on the repository's runbook config:
INT4 deployment (4ร H200, TP=4, 160K context):
vllm serve deepseek-ai/DeepSeek-R1 \
--tensor-parallel-size 4 \
--quantization bitsandbytes-nf4 \
--max-model-len 163840 \
--port 8000
FP8 deployment (6ร H200, TP=6, 160K context):
# Serve DeepSeek's FP8 weights (deepseek-ai/DeepSeek-R1 ships FP8 checkpoints;
# vLLM auto-detects quantization_config from config.json)
vllm serve deepseek-ai/DeepSeek-R1 \
--tensor-parallel-size 6 \
--pipeline-parallel-size 1 \
--quantization fp8 \
--max-model-len 163840 \
--port 8000 \
--gpu-memory-utilization 0.92
Key parameters explained
| Parameter | Value | Notes |
|---|---|---|
tensor-parallel-size | 4 (INT4) or 6 (FP8) | Do NOT use --trust-remote-code โ vLLM 0.6+ has native DeepSeek R1 support. TP=4 divides 64 routed experts evenly (16 per rank). |
max-model-len | 163840 | Supports full 160K context window (max_position_embeddings = 163840 in config.json) |
quantization (FP8) | fp8 | DeepSeek R1 ships native FP8 (e4m3) checkpoints; vLLM auto-detects from config.json, --quantization fp8 makes it explicit |
quantization (INT4) | bitsandbytes-nf4 | NF4 is the most accurate 4-bit format. Do NOT use bare bitsandbytes (old flag, defaults to INT8) |
gpu-memory-utilization | 0.92 (FP8) | Leave 8% headroom for OS and CUDA context. For INT4, vLLM auto-manages with default 0.90 |
vLLM compatibility: DeepSeek R1 is natively supported in vLLM 0.6+ (released September 2024). The --trust-remote-code flag is deprecated and unnecessary โ using it may produce warnings in newer vLLM versions. INT4 quantization via bitsandbytes-nf4 was tested with vLLM 0.6.2+.
Throughput expectations
Estimated output tokens per second with vLLM on H200 (based on DeepSeek API at 30 tok/s as reference, with local hardware advantage):
| Config | GPUs | Precision | Est. throughput | Source |
|---|---|---|---|---|
| TP=4, INT4 | 4x H200 | INT4 | 100โ150 tok/s | Extrapolated from model data |
| TP=6, FP8 | 6x H200 | FP8 | 120โ180 tok/s | Extrapolated from model data |
| TP=8, FP8 | 8x H200 | FP8 | 180โ250 tok/s | Even MoE split (8 experts per rank) |
Note: These are estimates. Actual throughput depends on context length, batch size, and MoE routing patterns. For production use, benchmark with your actual workload.
Decision table: choose your deployment
| Your situation | Recommendation | Why |
|---|---|---|
| <50M tokens/day, need flexibility | DeepSeek API ($0.55/M in) | Cheapest, no hardware investment |
| <305M tokens/day, need cost savings | NVIDIA NIM ($0.10/M in) | API is 230ร cheaper than self-hosting at this tier |
| >305M tokens/day, need privacy/compliance | 4x H200 INT4 (TP=4) | $11.16/hr breaks even at ~305M tokens/day |
| >305M tokens/day, 160K context needed | 6x H200 FP8 (TP=6) | Full context + even MoE split |
| Need <20ms latency, privacy required | 4x H200 INT4 (TP=4) | Even MoE distribution, predictable performance |
| Testing, prototyping only | GitHub Models (TRIAL) | Free tier with 15 RPM, 150 req/day |
| <100K tokens/day | API | Self-hosting never breaks even |
When API is always better
If you process fewer than ~305M tokens per day and can tolerate network latency, the API is always cheaper. At $0.878/M (DeepSeek) or $0.11/M (NVIDIA NIM), the API cost per token is orders of magnitude below self-hosting cost ($15.50โ38.76/M at 80โ200 tok/s).
When self-hosting wins
Self-hosting becomes cost-effective at volumes above ~305M tokens/day (4ร H200, INT4) primarily because:
- Fixed cost amortization: $267.84/day (4ร H200 at $2.79/hr) spread over sufficient tokens
- Throughput scales with GPU count: 4ร H200 can sustain 100โ200 tok/s, lowering per-token cost
- No rate limits: APIs cap at 30 tok/s per request (DeepSeek), a hard ceiling of ~2.6M tokens/day per key
For context: a single DeepSeek API key at 30 tok/s can process ~2.6M tokens/day. To match 4ร H200 at 120 tok/s (10.4M tokens/day), you'd need multiple API keys โ and under 24/7 rental you'd still only beat the API above ~305M tokens/day.
However, self-hosting requires:
- Kubernetes or equivalent orchestration
- Model monitoring and scaling
- GPU instance management
- Network and storage setup
- Ongoing maintenance
For most users, the API is the right choice unless you have dedicated DevOps resources and processing volumes exceeding 305M tokens/day.
Practical example: 70B-codebase summarization
Scenario: You have 5,000 code files averaging 500 lines each. You want to summarize each file using DeepSeek R1.
| Metric | Value |
|---|---|
| Total tokens | ~2.5M (input) |
| Output tokens | ~500K (summary per file) |
| Total tokens processed | 3M |
| API cost (DeepSeek) | 2.5M ร $0.55 + 500K ร $2.19 = $1.38 + $1.10 = $2.48 |
| API cost (NVIDIA NIM) | 2.5M ร $0.10 + 500K ร $0.15 = $0.25 + $0.08 = $0.33 |
| Self-hosted (4x H200, INT4, 10 min) | 4 ร $2.79 ร (10/60) = $1.86 |
In this scenario, NVIDIA NIM is cheapest, followed by DeepSeek API, then self-hosting. If you need to process 100 batches per day (300M tokens/day):
| Option | Daily cost | Notes |
|---|---|---|
| NVIDIA NIM | $0.33 ร 100 = $33 | Cheapest at any volume |
| DeepSeek API | $2.48 ร 100 = $248 | Reasonable up to 305M tokens/day |
| Self-hosted (4x H200) | $267.84 | Beats DeepSeek API at 305M+ tokens/day |
Limitations and caveats
-
GPU prices change daily. See cloud GPU pricing for the latest rates. This article's numbers are from the provider rate snapshot verified 2026-10-10.
-
Throughput estimates are extrapolated from the DeepSeek API rate of 30 tok/s. Local H200 throughput with vLLM + TP=4 will be higher (100โ200 tok/s), but exact numbers depend on your workload, batch size, and context length.
-
DeepSeek R1 does NOT require
--trust-remote-codein vLLM 0.6+. The model is natively supported. For INT4, use--quantization bitsandbytes-nf4(not the older barebitsandbytesflag, which defaults to INT8). -
INT4 weights (336 GB) + 4.19 GB MLA KV cache + 1.6 GB overhead = 342 GB subtotal. With 18% fragmentation headroom, total is ~403 GB. This fits on 3ร H200 (423 GB nominal), but 4ร H200 (TP=4) is recommended for even MoE expert distribution (64 experts รท 4 ranks = 16).
-
FP8 weights (671 GB) + 4.19 GB KV cache + 1.6 GB overhead = 676.8 GB subtotal. With 18% fragmentation headroom, total is ~799 GB. This requires 6ร H200 (846 GB nominal) at TP=6, or 8ร H200 (1,128 GB) at TP=8 for even expert distribution.
-
DeepSeek R1 API may be unavailable on DeepSeek's pricing page as of 2026-10-11. The $0.55/M rate reflects historical pricing and third-party provider rates. Verify current pricing at DeepSeek API docs.
-
NVIDIA NIM at $0.10/M input is the cheapest option in our registry table, but it is a third-party/registry-sourced rate (not re-verified against NVIDIA's live catalog for this article) and requires an NVIDIA Developer Program account. Verify current NIM pricing at nvidia.com before modeling costs. See NVIDIA NIM provider page.
-
The 4:1 input/output ratio (80% input, 20% output) is typical for RAG and code generation. Adjust the breakeven calculation if your ratio differs significantly.
-
MoE TP sizing: DeepSeek R1 has 64 routed experts. For even distribution, TP must divide 64 evenly: TP=4 (16 experts/rank) or TP=8 (8 experts/rank). TP=6 (10.67 per rank) is acceptable but slightly imbalanced.
-
Fragmentation rate: The calculator engine uses 18% as a conservative estimate; vLLM's actual allocator typically achieves ~10%. Both yield the same GPU count recommendations.
Methodology
- Architecture data: Verified from HuggingFace config.json (num_hidden_layers=61, num_attention_heads=128, num_key_value_heads=128, max_position_embeddings=163840, q_lora_dim=512, qk_rope_head_dim=64, n_routed_experts=64, n_shared_expert=2, n_activated_experts=2). Verified 2026-10-11.
- VRAM figures: From vram-calculator.ts with 18% fragmentation (conservative). Weights: FP16 = params ร 2 bytes, FP8/INT8 = params ร 1 byte, INT4 = params ร 0.5 bytes. KV cache uses MLA formula:
61 ร (512+64) ร context ร 1 bytefor FP8/INT4/INT8, ร 2 bytes for FP16. - API pricing: From models-registry.json:apiProviders, verified 2026-10-11. DeepSeek R1 no longer listed on official pricing page โ pricing reflects third-party provider rates.
- Cloud GPU rates: From H200 cloud pricing page, daily snapshot 2026-10-10.
- Benchmarks: From models-registry.json:benchmarks. Verified from respective benchmark platforms (SWE-bench, LiveCodeBench, AIME 2024, MMLU).
- vLLM config: Based on vLLM 0.6+ native DeepSeek R1 support. No
--trust-remote-codeneeded. - Throughput estimates: Extrapolated from DeepSeek API's documented 30 tok/s rate (models-registry.json) and typical multi-GPU H200 performance with vLLM.
- Breakeven calculation: Daily self-hosted cost รท API cost per million tokens. Assumes 80% input / 20% output token mix, 24/7 operation, and Vast.ai H200 spot rates.
Related resources
- DeepSeek R1 model page โ full specs, VRAM table, API pricing
- NVIDIA H200 cloud pricing โ live H200 rates across providers
- NVIDIA B200 cloud pricing โ B200 rates and specs
- LLM VRAM calculator โ calculate VRAM for any model and GPU
- H200 vs B200 GPU comparison โ throughput and specs side-by-side
- DeepSeek R1: FP8 vs INT4 memory analysis โ detailed quantization breakdown
- How much does it cost to run DeepSeek R1? โ cost analysis overview
- DeepSeek R1 company page โ API configuration and free tiers
Related Articles
How to Fine-Tune an LLM with Limited GPU Memory in 2026
Fine-tune an LLM on limited GPU memory: full fine-tune vs LoRA vs QLoRA memory math, fit tables for 12-80GB GPUs, and checkpointing tricks.
InfrastructureWhich LLMs Can Run on 8GB, 16GB, 24GB, 48GB and 80GB GPUs in 2026?
Which LLMs fit on 8GB, 16GB, 24GB, 48GB and 80GB GPUs โ verified fit tables by quantization and context from OpenGPU Radar's VRAM engine.
InfrastructureB200 NVLink 5.0 Scaling: When Does 8x B200 Beat 16x H100?
8x B200 SXM (NVLink 5.0 at 1.8 TB/s) equals 16x H100's aggregate NVLink bandwidth at similar cost โ and delivers 50% more 70B inference throughput. Here's the math on bandwidth, power, and cost.
InfrastructureLLM VRAM Calculator: GPU Requirements for Llama, Qwen, DeepSeek & Claude in 2026
Find the right GPU for any LLM: Llama 3.3 70B needs 71 GB of FP8 weights (~81 GB in service), DeepSeek R1 671B needs 786 GB at 128K context. VRAM sizing chart by model.