DeepSeek V3 VRAM Requirements: FP16 / INT4 / KV-Cache (2026)
DeepSeek V3 runs on 671B parameters with a 128K-token context window. Weights are deterministic param math; KV-cache uses the registry layer geometry where published, otherwise a modeled GQA estimate โ every figure is provenance-labeled below.
1340 GB FP16 weights
weights = parametersB ร bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)
336 GB INT4 weights
weights = parametersB ร bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)
128K context window
How much VRAM does DeepSeek V3 need?
Weights: 1340 GB at FP16, 336 GB at INT4 (parameters ร bytes-per-parameter: FP16 = 2 B, INT4 = 0.5 B).
Full stack at 128K: 1498.4 GB FP16 / 381.0 GB INT4 โ including KV-cache, CUDA overhead, activations, and fragmentation headroom from the canonical VRAM engine.
Single-GPU verdict: Multi-GPU required โ 381 GB INT4 exceeds single-GPU capacity; registry recommendation: 8x H200.
VRAM by Precision & Context
From the canonical VRAM engine: weights + KV-cache + CUDA overhead + activations + fragmentation headroom. KV-cache uses published layer geometry where available; otherwise a modeled GQA estimate (labeled in the badge).
| Precision | Context | Weights | KV-Cache | Total (est.) |
|---|---|---|---|---|
| FP16 | 4K | 1342.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 1,342 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.59 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.59 GB Methodology GQA KV Cache: kvBytes = 2 ร numLayers ร numKvHeads ร headDim ร contextLength ร batchSize ร bytesPerKvElement. GQA ratio: 4:1. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 1478.6 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 1478.61 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| FP16 | 33K | 1342.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 1,342 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 4.75 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 4.75 GB Methodology GQA KV Cache: kvBytes = 2 ร numLayers ร numKvHeads ร headDim ร contextLength ร batchSize ร bytesPerKvElement. GQA ratio: 4:1. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 1483.2 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 1483.18 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| FP16 | 128K | 1342.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 1,342 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 18.55 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 18.55 GB Methodology GQA KV Cache: kvBytes = 2 ร numLayers ร numKvHeads ร headDim ร contextLength ร batchSize ร bytesPerKvElement. GQA ratio: 4:1. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 1498.4 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 1498.37 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| INT4 | 4K | 335.5 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 335.5 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.30 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.3 GB Methodology GQA KV Cache: kvBytes = 2 ร numLayers ร numKvHeads ร headDim ร contextLength ร batchSize ร bytesPerKvElement. GQA ratio: 4:1. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 371.1 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 371.14 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| INT4 | 33K | 335.5 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 335.5 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 2.38 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 2.38 GB Methodology GQA KV Cache: kvBytes = 2 ร numLayers ร numKvHeads ร headDim ร contextLength ร batchSize ร bytesPerKvElement. GQA ratio: 4:1. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 373.4 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 373.42 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| INT4 | 128K | 335.5 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 335.5 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 9.28 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 9.28 GB Methodology GQA KV Cache: kvBytes = 2 ร numLayers ร numKvHeads ร headDim ร contextLength ร batchSize ร bytesPerKvElement. GQA ratio: 4:1. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 381.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 381.01 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
Cloud GPUs That Fit DeepSeek V3
Fit = full-stack total (weights + KV at 128K) โค GPU VRAM. Rates are the lowest observed on-demand rows in data/providers.json, refreshed daily.
| GPU | VRAM | FP16 fit | INT4 fit | Lowest on-demand |
|---|---|---|---|---|
| H200 SXM5 | 141 GB | OOM | OOM | $2.79/hr (Vast.ai) |
| B200 Blackwell | 192 GB | OOM | OOM | $3.99/hr (Vast.ai) |
| H100 SXM5 | 80 GB | OOM | OOM | $1.89/hr (Vast.ai) |
| A100 80GB SXM4 | 80 GB | OOM | OOM | $1.59/hr (Lambda Labs) |
| L40S | 48 GB | OOM | OOM | $0.69/hr (Vast.ai) |
| RTX 4090 | 24 GB | OOM | OOM | $0.34/hr (Vast.ai) |
Recommended GPUs for DeepSeek V3
Minimum = smallest single GPU that fits the full INT4 stack (381.0 GB). Optimal = registry recommendation, falling back to the lowest observed on-demand rate among fitting cards. Every card links to its canonical GPU page.
Workload Scenarios Using DeepSeek V3
Costed workload guides that price DeepSeek V3 against verified cloud rates and registry VRAM math.
Next steps