Infrastructure2026-10-11โ€ขBy Sreeโ€ข5 min read

DeepSeek R1 Serving Guide: FP8 vs INT4 VRAM, vLLM Config, and H200/B200 Cost Analysis

DeepSeek R1 671B MoE serving guide: INT4 = 403 GB VRAM (4x H200, TP=4); FP8 = 799 GB (6x H200). vLLM configuration with bitsandbytes-nf4 quantization and breakeven analysis at $0.55/M API vs $11.16/hr self-hosted.

Direct answer

DeepSeek R1 (671B MoE, 37B active) fits on 4ร— H200 with INT4 quantized weights (โ‰ˆ375 GB VRAM at 128K context with 10% fragmentation headroom, TP=4 recommended for even MoE expert split; 3ร— H200 is the hard minimum at 423 GB) or 6ร— H200 with FP8 (โ‰ˆ745 GB, TP=6). Using a more conservative 18% allocator headroom (the /calculator setting) yields โ‰ˆ403 GB INT4 / โ‰ˆ799 GB FP8 โ€” same GPU counts. For most workloads, an R1-compatible API is cheaper than self-hosting; self-hosting only becomes cost-competitive above ~305M tokens/day under a 24/7 rental assumption. Use vLLM 0.6+ with --tensor-parallel-size 4 (INT4) or --tensor-parallel-size 6 (FP8) and --max-model-len 163840 for full 160K context.


Why DeepSeek R1 deployment is different

DeepSeek R1 is a Mixture-of-Experts (MoE) model with 671B total parameters but only 37B active at inference time (5.5% of total). This means:

MetricValueSource
Total parameters671Bmodels-registry.json:deepseek-r1
Active parameters (per token)37Bmodels-registry.json:deepseek-r1
ArchitectureMoE + MLAconfig.json
Context window160,000 tokensconfig.json max_position_embeddings=163840
Layers61config.json num_hidden_layers=61
Attention heads128config.json num_attention_heads=128
KV heads128config.json num_key_value_heads=128
Head dimension128config.json
Routed experts64config.json n_routed_experts=64
kvLoraRank512config.json kv_lora_rank=512
qkRopeHeadDim64config.json qk_rope_head_dim=64

Unlike dense models, the 671B parameter count can be misleading โ€” only 37B are active per forward pass. However, all 671B weights must be loaded into memory (either VRAM or CPU/disk with offloading) because the model dynamically routes tokens to different experts.

DeepSeek R1 uses Multi-Latent Attention (MLA) rather than traditional Grouped-Query Attention (GQA). This compresses the KV cache significantly: instead of storing per-head keys and values (128 heads ร— 128 dims ร— 2 = 32,768 values/token/layer), MLA stores a compressed latent of kvLoraRank (512) + qkRopeHeadDim (64) = 576 dimensions per token per layer โ€” a 57ร— reduction in KV cache memory.

Benchmarks (from models-registry.json):

BenchmarkDeepSeek R1DeepSeek R1 Distill 70BGPT-4o
SweBench (coding)49.250.849.0
LiveCodeBench65.960.142.0
AIME 2024 (math)79.884.675.5
Math50097.392.575.5
MMLU (general knowledge)90.888.086.4

R1 leads on math and coding benchmarks among open-weight models, which justifies the deployment complexity for reasoning-heavy workloads.


VRAM requirements by quantization level

All VRAM numbers include KV cache for 128K context, MLA KV cache for 160K is 5.36 GB (FP8/INT4/INT8) or 10.72 GB (FP16). Sources: models-registry.json, config.json, and deepseek-r1-fp8-vs-int4-memory. The calculator engine at /calculator uses 18% fragmentation as a conservative estimate; the canonical engine uses 10% (matching vLLM's allocator behavior).

Memory breakdown components

ComponentDescriptionValue (DeepSeek R1)
Model weightsStatic parameters loaded in VRAM. FP16 = 2 bytes/param, FP8 = 1 byte/param, INT4 = 0.5 bytes/param. MoE means 671B total but only 37B active per token.FP8: 671 GB, INT4: 336 GB, FP16: 1,342 GB
KV cacheAttention context storage per request. MLA compresses to kvLoraRank (512) + qkRopeHeadDim (64) = 576 dims per token per layer. Stored at 1 byte/element for FP8/INT4/INT8, 2 bytes for FP16.FP8/INT4: 4.19 GB @ 128K, FP16: 8.38 GB @ 128K
CUDA runtime overheadFixed PyTorch/vLLM CUDA context, memory allocator buffers.1.2 GB (fixed)
Activation memoryAutoregressive decode workspace for batch=1.0.4 GB (fixed)
Fragmentation headroom10โ€“18% safety margin for memory allocator fragmentation.10% (canonical) or 18% (calculator)

VRAM totals at 128K context

QuantizationWeightsKV Cache (MLA)CUDAActivationsFragmentationTotal VRAM
FP161,342 GB8.38 GB1.2 GB0.4 GB136.0 GB (10%)1,487.2 GB
FP8671 GB4.19 GB1.2 GB0.4 GB68.0 GB (10%)744.5 GB
INT4335.5 GB4.19 GB1.2 GB0.4 GB34.3 GB (10%)375.4 GB
INT8671 GB4.19 GB1.2 GB0.4 GB68.0 GB (10%)744.5 GB

Note: FP8 and INT8 produce identical VRAM totals (both use 1 byte/param for weights). INT4 is the only quantization that meaningfully reduces VRAM. KV cache is precision-invariant for FP8/INT4/INT8 (1 byte/element) and doubles for FP16 (2 bytes/element).

How many GPUs do you need? (128K context, 18% fragmentation)

GPUVRAM (each)FP8 (799 GB)INT4 (403 GB)INT8 (799 GB)
H200 SXM5 (141 GB)141 GB6 GPUs3 GPUs6 GPUs
B200 (192 GB)192 GB5 GPUs3 GPUs5 GPUs
H100 SXM5 (80 GB)80 GB10 GPUs6 GPUs10 GPUs
RTX 4090 (24 GB)24 GBโ€”โ€”โ€”

Recommendation: Use 4ร— H200 with TP=4 for INT4 serving. This provides even MoE expert distribution (64 experts รท 4 ranks = 16 per rank) and 161 GB headroom over the 403 GB requirement. For FP8, use 6ร— H200 with TP=6 โ€” sufficient VRAM (846 GB vs 799 GB needed). For even MoE distribution with FP8, TP=8 provides 64 รท 8 = 8 experts per rank, at the cost of 2 additional GPUs.


Cost analysis: API vs self-hosted

DeepSeek API pricing

ProviderInputOutputThroughputSource
DeepSeek$0.55/M$2.19/M30 tok/smodels-registry.json
NVIDIA NIM$0.10/M$0.15/M500 tok/smodels-registry.json
GitHub Models$0.15/M$0.20/M30 tok/smodels-registry.json
OpenRouter$0.70/M$2.50/M22 tok/smodels-registry.json
Together AI$0.80/M$2.40/M25 tok/smodels-registry.json
Fireworks$0.90/M$2.50/M28 tok/smodels-registry.json

Note: As of 2026-10-11, DeepSeek R1 is no longer listed on DeepSeek's pricing page. The $0.55/M rate reflects historical DeepSeek API pricing; Together, OpenRouter, Fireworks, GitHub Models, and NIM rates are registry-sourced third-party prices. Verify current pricing with each provider before deployment โ€” do not treat this table as a live price feed.

Self-hosting costs (cloud GPU rates)

GPUProviderHourlyMonthly (720h)
H200 SXMVast.ai$2.79$1,707.20
H200 SXMSpheron$3.19$1,952.00
H200 SXMRunPod$4.31$2,638.40
B200 SXMVast.ai$3.99$2,441.60
B200 SXMSpheron$4.49$2,748.80
H100 SXMVast.ai$1.89$1,076.80

All provider rates from H200 cloud pricing, verified 2026-10-10.

Breakeven calculation

Assumptions:

  • 4ร— H200 INT4 (TP=4) on Vast.ai spot: 4 ร— $2.79 = $11.16/hour
  • 24-hour continuous operation: $11.16 ร— 24 = $267.84 per day
  • Local throughput: 120 tokens/sec (realistic for TP=4 with INT4, vLLM 0.6+)
  • Daily throughput: 120 ร— 3,600 ร— 24 = 10.37M tokens/day
  • API token mix: 80% input, 20% output (standard 4:1 ratio for RAG/codegen)

API average cost (DeepSeek, 4:1 ratio):

  • $0.55 ร— 0.8 + $2.19 ร— 0.2 = $0.878 per million tokens

Self-hosted cost per million tokens (4ร— H200, INT4, 120 tok/s):

  • $11.16 / (120 ร— 3,600 / 1,000,000) = $25.84 per million tokens

Breakeven (24/7 rental assumption): self-hosted daily cost is a fixed $267.84 whether you use 1M tokens or 500M โ€” you pay for idle GPUs. API cost scales linearly with volume at $0.878/M. Self-hosting becomes cheaper when:

daily_tokens ร— $0.878 / 1,000,000 > $267.84
daily_tokens > ~305M
Daily token volumeAPI cost (DeepSeek, 4:1 mix)Self-hosted 4ร— H200 (24/7)Cheaper option
1M tokens$0.88$267.84API
10M tokens$8.78$267.84API
50M tokens$43.90$267.84API
100M tokens$87.80$267.84API
200M tokens$175.60$267.84API
305M tokens$267.79$267.84breakeven
500M tokens$439.00$267.84Self-hosted

If you rent GPUs only for the hours you actually generate (no idle time), self-hosted cost is $25.84 per million tokens at 120 tok/s โ€” still far above the API's $0.878/M, so API wins at every volume. The 305M/day breakeven exists only under 24/7 committed rental (or reserved instances) where idle capacity is sunk cost. Pick the framing that matches how you'd actually procure GPUs.

Breakeven summary: Self-hosting on 4ร— H200 (INT4, TP=4) at 120 tok/s becomes cheaper than the DeepSeek API above ~305M tokens/day โ€” but only if the cluster runs 24/7. Under on-demand, pay-for-what-you-use renting, the API is always cheaper per token. Self-hosting also wins earlier when you need privacy, deterministic latency, or freedom from API rate limits (see below).

ThroughputTokens/hrDaily tokensDaily self-hosted costCost per M tokens (self-hosted)Breakeven daily volume
80 tok/s288K6.91M$267.84$38.76305M
100 tok/s360K8.64M$267.84$31.00305M
120 tok/s432K10.37M$267.84$25.84305M
150 tok/s540K12.96M$267.84$20.67305M
200 tok/s720K17.28M$267.84$15.50305M

Note: The breakeven daily volume is the same regardless of throughput โ€” it's driven by the fixed daily rental cost ($267.84) and the API cost per million tokens ($0.878/M). Throughput only affects the per-token cost of self-hosting.

Key insight: DeepSeek's API is extremely cost-competitive ($0.878/M) compared to most providers. Only NVIDIA NIM ($0.11/M) and GitHub Models ($0.15/M) are cheaper per-token. Self-hosting only wins at very high daily volumes (>305M tokens) or when you need:

  1. Deterministic latency (no network round-trip)
  2. Data privacy (can't send prompts to external APIs)
  3. Unlimited rate limits (API caps at 30 tokens/sec โ†’ max ~2.6M tokens/day per key)
  4. Custom model modifications (e.g., custom loss functions, LoRA adapters)

For comparison with other models:

ModelAPI inputAPI outputSelf-hosted est.Notes
DeepSeek R1$0.55/M$2.19/M$15.50โ€“38.76/M (4x H200, 80โ€“200 tok/s)API wins unless >305M tokens/day
Llama 3.3 70B$0.88/M$0.88/M~$12/M (1x H200)Simpler deployment, lower VRAM
Qwen 2.5 72B$0.20/M$0.60/M~$12/M (1x H200)Better cost-to-quality ratio

See DeepSeek R1 model page for full API provider comparisons and cloud GPU pricing for live rate data.


vLLM configuration

Based on the repository's runbook config:

INT4 deployment (4ร— H200, TP=4, 160K context):

vllm serve deepseek-ai/DeepSeek-R1 \
  --tensor-parallel-size 4 \
  --quantization bitsandbytes-nf4 \
  --max-model-len 163840 \
  --port 8000

FP8 deployment (6ร— H200, TP=6, 160K context):

# Serve DeepSeek's FP8 weights (deepseek-ai/DeepSeek-R1 ships FP8 checkpoints;
# vLLM auto-detects quantization_config from config.json)
vllm serve deepseek-ai/DeepSeek-R1 \
  --tensor-parallel-size 6 \
  --pipeline-parallel-size 1 \
  --quantization fp8 \
  --max-model-len 163840 \
  --port 8000 \
  --gpu-memory-utilization 0.92

Key parameters explained

ParameterValueNotes
tensor-parallel-size4 (INT4) or 6 (FP8)Do NOT use --trust-remote-code โ€” vLLM 0.6+ has native DeepSeek R1 support. TP=4 divides 64 routed experts evenly (16 per rank).
max-model-len163840Supports full 160K context window (max_position_embeddings = 163840 in config.json)
quantization (FP8)fp8DeepSeek R1 ships native FP8 (e4m3) checkpoints; vLLM auto-detects from config.json, --quantization fp8 makes it explicit
quantization (INT4)bitsandbytes-nf4NF4 is the most accurate 4-bit format. Do NOT use bare bitsandbytes (old flag, defaults to INT8)
gpu-memory-utilization0.92 (FP8)Leave 8% headroom for OS and CUDA context. For INT4, vLLM auto-manages with default 0.90

vLLM compatibility: DeepSeek R1 is natively supported in vLLM 0.6+ (released September 2024). The --trust-remote-code flag is deprecated and unnecessary โ€” using it may produce warnings in newer vLLM versions. INT4 quantization via bitsandbytes-nf4 was tested with vLLM 0.6.2+.

Throughput expectations

Estimated output tokens per second with vLLM on H200 (based on DeepSeek API at 30 tok/s as reference, with local hardware advantage):

ConfigGPUsPrecisionEst. throughputSource
TP=4, INT44x H200INT4100โ€“150 tok/sExtrapolated from model data
TP=6, FP86x H200FP8120โ€“180 tok/sExtrapolated from model data
TP=8, FP88x H200FP8180โ€“250 tok/sEven MoE split (8 experts per rank)

Note: These are estimates. Actual throughput depends on context length, batch size, and MoE routing patterns. For production use, benchmark with your actual workload.


Decision table: choose your deployment

Your situationRecommendationWhy
<50M tokens/day, need flexibilityDeepSeek API ($0.55/M in)Cheapest, no hardware investment
<305M tokens/day, need cost savingsNVIDIA NIM ($0.10/M in)API is 230ร— cheaper than self-hosting at this tier
>305M tokens/day, need privacy/compliance4x H200 INT4 (TP=4)$11.16/hr breaks even at ~305M tokens/day
>305M tokens/day, 160K context needed6x H200 FP8 (TP=6)Full context + even MoE split
Need <20ms latency, privacy required4x H200 INT4 (TP=4)Even MoE distribution, predictable performance
Testing, prototyping onlyGitHub Models (TRIAL)Free tier with 15 RPM, 150 req/day
<100K tokens/dayAPISelf-hosting never breaks even

When API is always better

If you process fewer than ~305M tokens per day and can tolerate network latency, the API is always cheaper. At $0.878/M (DeepSeek) or $0.11/M (NVIDIA NIM), the API cost per token is orders of magnitude below self-hosting cost ($15.50โ€“38.76/M at 80โ€“200 tok/s).

When self-hosting wins

Self-hosting becomes cost-effective at volumes above ~305M tokens/day (4ร— H200, INT4) primarily because:

  1. Fixed cost amortization: $267.84/day (4ร— H200 at $2.79/hr) spread over sufficient tokens
  2. Throughput scales with GPU count: 4ร— H200 can sustain 100โ€“200 tok/s, lowering per-token cost
  3. No rate limits: APIs cap at 30 tok/s per request (DeepSeek), a hard ceiling of ~2.6M tokens/day per key

For context: a single DeepSeek API key at 30 tok/s can process ~2.6M tokens/day. To match 4ร— H200 at 120 tok/s (10.4M tokens/day), you'd need multiple API keys โ€” and under 24/7 rental you'd still only beat the API above ~305M tokens/day.

However, self-hosting requires:

  • Kubernetes or equivalent orchestration
  • Model monitoring and scaling
  • GPU instance management
  • Network and storage setup
  • Ongoing maintenance

For most users, the API is the right choice unless you have dedicated DevOps resources and processing volumes exceeding 305M tokens/day.


Practical example: 70B-codebase summarization

Scenario: You have 5,000 code files averaging 500 lines each. You want to summarize each file using DeepSeek R1.

MetricValue
Total tokens~2.5M (input)
Output tokens~500K (summary per file)
Total tokens processed3M
API cost (DeepSeek)2.5M ร— $0.55 + 500K ร— $2.19 = $1.38 + $1.10 = $2.48
API cost (NVIDIA NIM)2.5M ร— $0.10 + 500K ร— $0.15 = $0.25 + $0.08 = $0.33
Self-hosted (4x H200, INT4, 10 min)4 ร— $2.79 ร— (10/60) = $1.86

In this scenario, NVIDIA NIM is cheapest, followed by DeepSeek API, then self-hosting. If you need to process 100 batches per day (300M tokens/day):

OptionDaily costNotes
NVIDIA NIM$0.33 ร— 100 = $33Cheapest at any volume
DeepSeek API$2.48 ร— 100 = $248Reasonable up to 305M tokens/day
Self-hosted (4x H200)$267.84Beats DeepSeek API at 305M+ tokens/day

Limitations and caveats

  1. GPU prices change daily. See cloud GPU pricing for the latest rates. This article's numbers are from the provider rate snapshot verified 2026-10-10.

  2. Throughput estimates are extrapolated from the DeepSeek API rate of 30 tok/s. Local H200 throughput with vLLM + TP=4 will be higher (100โ€“200 tok/s), but exact numbers depend on your workload, batch size, and context length.

  3. DeepSeek R1 does NOT require --trust-remote-code in vLLM 0.6+. The model is natively supported. For INT4, use --quantization bitsandbytes-nf4 (not the older bare bitsandbytes flag, which defaults to INT8).

  4. INT4 weights (336 GB) + 4.19 GB MLA KV cache + 1.6 GB overhead = 342 GB subtotal. With 18% fragmentation headroom, total is ~403 GB. This fits on 3ร— H200 (423 GB nominal), but 4ร— H200 (TP=4) is recommended for even MoE expert distribution (64 experts รท 4 ranks = 16).

  5. FP8 weights (671 GB) + 4.19 GB KV cache + 1.6 GB overhead = 676.8 GB subtotal. With 18% fragmentation headroom, total is ~799 GB. This requires 6ร— H200 (846 GB nominal) at TP=6, or 8ร— H200 (1,128 GB) at TP=8 for even expert distribution.

  6. DeepSeek R1 API may be unavailable on DeepSeek's pricing page as of 2026-10-11. The $0.55/M rate reflects historical pricing and third-party provider rates. Verify current pricing at DeepSeek API docs.

  7. NVIDIA NIM at $0.10/M input is the cheapest option in our registry table, but it is a third-party/registry-sourced rate (not re-verified against NVIDIA's live catalog for this article) and requires an NVIDIA Developer Program account. Verify current NIM pricing at nvidia.com before modeling costs. See NVIDIA NIM provider page.

  8. The 4:1 input/output ratio (80% input, 20% output) is typical for RAG and code generation. Adjust the breakeven calculation if your ratio differs significantly.

  9. MoE TP sizing: DeepSeek R1 has 64 routed experts. For even distribution, TP must divide 64 evenly: TP=4 (16 experts/rank) or TP=8 (8 experts/rank). TP=6 (10.67 per rank) is acceptable but slightly imbalanced.

  10. Fragmentation rate: The calculator engine uses 18% as a conservative estimate; vLLM's actual allocator typically achieves ~10%. Both yield the same GPU count recommendations.


Methodology

  • Architecture data: Verified from HuggingFace config.json (num_hidden_layers=61, num_attention_heads=128, num_key_value_heads=128, max_position_embeddings=163840, q_lora_dim=512, qk_rope_head_dim=64, n_routed_experts=64, n_shared_expert=2, n_activated_experts=2). Verified 2026-10-11.
  • VRAM figures: From vram-calculator.ts with 18% fragmentation (conservative). Weights: FP16 = params ร— 2 bytes, FP8/INT8 = params ร— 1 byte, INT4 = params ร— 0.5 bytes. KV cache uses MLA formula: 61 ร— (512+64) ร— context ร— 1 byte for FP8/INT4/INT8, ร— 2 bytes for FP16.
  • API pricing: From models-registry.json:apiProviders, verified 2026-10-11. DeepSeek R1 no longer listed on official pricing page โ€” pricing reflects third-party provider rates.
  • Cloud GPU rates: From H200 cloud pricing page, daily snapshot 2026-10-10.
  • Benchmarks: From models-registry.json:benchmarks. Verified from respective benchmark platforms (SWE-bench, LiveCodeBench, AIME 2024, MMLU).
  • vLLM config: Based on vLLM 0.6+ native DeepSeek R1 support. No --trust-remote-code needed.
  • Throughput estimates: Extrapolated from DeepSeek API's documented 30 tok/s rate (models-registry.json) and typical multi-GPU H200 performance with vLLM.
  • Breakeven calculation: Daily self-hosted cost รท API cost per million tokens. Assumes 80% input / 20% output token mix, 24/7 operation, and Vast.ai H200 spot rates.

Related resources