VRAM Reference

DeepSeek V3 VRAM Requirements: FP16 / INT4 / KV-Cache (2026)

DeepSeek V3 runs on 671B parameters with a 128K-token context window. Weights are deterministic param math; KV-cache uses the registry layer geometry where published, otherwise a modeled GQA estimate โ€” every figure is provenance-labeled below.

FP16 weights1340 GB
INT4 weights336 GB
Context128K tokens
Single-GPU verdictMulti-GPU (381 GB)
1340 GB FP16 weights
Calculated EstimateHIGH
SourceOpenGPU Radar models-registry.json (param-count weight model)
VerifiedSep 29, 2026
Value
1,340 GB
Methodology

weights = parametersB ร— bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)

Refreshed daily from provider APIs and market scrapingSep 29, 2026
336 GB INT4 weights
Calculated EstimateHIGH
SourceOpenGPU Radar models-registry.json (param-count weight model)
VerifiedSep 29, 2026
Value
336 GB
Methodology

weights = parametersB ร— bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)

Refreshed daily from provider APIs and market scrapingSep 29, 2026
128K context window
Manufacturer SpecHIGH
SourceModel vendor context-window specification
VerifiedSep 29, 2026
Value
128,000 tokens
Refreshed daily from provider APIs and market scrapingSep 29, 2026
Methodology โ†’

How much VRAM does DeepSeek V3 need?

Weights: 1340 GB at FP16, 336 GB at INT4 (parameters ร— bytes-per-parameter: FP16 = 2 B, INT4 = 0.5 B).

Full stack at 128K: 1498.4 GB FP16 / 381.0 GB INT4 โ€” including KV-cache, CUDA overhead, activations, and fragmentation headroom from the canonical VRAM engine.

Single-GPU verdict: Multi-GPU required โ€” 381 GB INT4 exceeds single-GPU capacity; registry recommendation: 8x H200.

VRAM by Precision & Context

From the canonical VRAM engine: weights + KV-cache + CUDA overhead + activations + fragmentation headroom. KV-cache uses published layer geometry where available; otherwise a modeled GQA estimate (labeled in the badge).

PrecisionContextWeightsKV-CacheTotal (est.)
FP164K1342.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
1,342 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: fp16
  • bytesPerParam: 2
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.59 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.59 GB
Methodology

GQA KV Cache: kvBytes = 2 ร— numLayers ร— numKvHeads ร— headDim ร— contextLength ร— batchSize ร— bytesPerKvElement. GQA ratio: 4:1.

Assumptions & Parameters
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • contextLength: 4096
  • batchSize: 1
  • quantization: fp16
  • gqaRatio: 4
Refreshed daily from provider APIs and market scrapingSep 26, 2026
1478.6 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
1478.61 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: fp16
  • contextLength: 4096
  • batchSize: 1
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • gqaRatio: 4
  • architecture: deepseek-v3
  • bytesPerParam: 2
  • bytesPerKvElement: 2
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: true
Refreshed daily from provider APIs and market scrapingSep 26, 2026
FP1633K1342.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
1,342 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: fp16
  • bytesPerParam: 2
Refreshed daily from provider APIs and market scrapingSep 26, 2026
4.75 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
4.75 GB
Methodology

GQA KV Cache: kvBytes = 2 ร— numLayers ร— numKvHeads ร— headDim ร— contextLength ร— batchSize ร— bytesPerKvElement. GQA ratio: 4:1.

Assumptions & Parameters
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • contextLength: 32768
  • batchSize: 1
  • quantization: fp16
  • gqaRatio: 4
Refreshed daily from provider APIs and market scrapingSep 26, 2026
1483.2 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
1483.18 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: fp16
  • contextLength: 32768
  • batchSize: 1
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • gqaRatio: 4
  • architecture: deepseek-v3
  • bytesPerParam: 2
  • bytesPerKvElement: 2
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: true
Refreshed daily from provider APIs and market scrapingSep 26, 2026
FP16128K1342.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
1,342 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: fp16
  • bytesPerParam: 2
Refreshed daily from provider APIs and market scrapingSep 26, 2026
18.55 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
18.55 GB
Methodology

GQA KV Cache: kvBytes = 2 ร— numLayers ร— numKvHeads ร— headDim ร— contextLength ร— batchSize ร— bytesPerKvElement. GQA ratio: 4:1.

Assumptions & Parameters
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • contextLength: 128000
  • batchSize: 1
  • quantization: fp16
  • gqaRatio: 4
Refreshed daily from provider APIs and market scrapingSep 26, 2026
1498.4 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
1498.37 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: fp16
  • contextLength: 128000
  • batchSize: 1
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • gqaRatio: 4
  • architecture: deepseek-v3
  • bytesPerParam: 2
  • bytesPerKvElement: 2
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: true
Refreshed daily from provider APIs and market scrapingSep 26, 2026
INT44K335.5 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
335.5 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: int4
  • bytesPerParam: 0.5
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.30 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.3 GB
Methodology

GQA KV Cache: kvBytes = 2 ร— numLayers ร— numKvHeads ร— headDim ร— contextLength ร— batchSize ร— bytesPerKvElement. GQA ratio: 4:1.

Assumptions & Parameters
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • contextLength: 4096
  • batchSize: 1
  • quantization: int4
  • gqaRatio: 4
Refreshed daily from provider APIs and market scrapingSep 26, 2026
371.1 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
371.14 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: int4
  • contextLength: 4096
  • batchSize: 1
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • gqaRatio: 4
  • architecture: deepseek-v3
  • bytesPerParam: 0.5
  • bytesPerKvElement: 1
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: true
Refreshed daily from provider APIs and market scrapingSep 26, 2026
INT433K335.5 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
335.5 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: int4
  • bytesPerParam: 0.5
Refreshed daily from provider APIs and market scrapingSep 26, 2026
2.38 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
2.38 GB
Methodology

GQA KV Cache: kvBytes = 2 ร— numLayers ร— numKvHeads ร— headDim ร— contextLength ร— batchSize ร— bytesPerKvElement. GQA ratio: 4:1.

Assumptions & Parameters
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • contextLength: 32768
  • batchSize: 1
  • quantization: int4
  • gqaRatio: 4
Refreshed daily from provider APIs and market scrapingSep 26, 2026
373.4 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
373.42 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: int4
  • contextLength: 32768
  • batchSize: 1
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • gqaRatio: 4
  • architecture: deepseek-v3
  • bytesPerParam: 0.5
  • bytesPerKvElement: 1
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: true
Refreshed daily from provider APIs and market scrapingSep 26, 2026
INT4128K335.5 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
335.5 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: int4
  • bytesPerParam: 0.5
Refreshed daily from provider APIs and market scrapingSep 26, 2026
9.28 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
9.28 GB
Methodology

GQA KV Cache: kvBytes = 2 ร— numLayers ร— numKvHeads ร— headDim ร— contextLength ร— batchSize ร— bytesPerKvElement. GQA ratio: 4:1.

Assumptions & Parameters
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • contextLength: 128000
  • batchSize: 1
  • quantization: int4
  • gqaRatio: 4
Refreshed daily from provider APIs and market scrapingSep 26, 2026
381.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
381.01 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 671
  • quantization: int4
  • contextLength: 128000
  • batchSize: 1
  • numLayers: 38
  • numKvHeads: 8
  • headDim: 128
  • gqaRatio: 4
  • architecture: deepseek-v3
  • bytesPerParam: 0.5
  • bytesPerKvElement: 1
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: true
Refreshed daily from provider APIs and market scrapingSep 26, 2026

Cloud GPUs That Fit DeepSeek V3

Fit = full-stack total (weights + KV at 128K) โ‰ค GPU VRAM. Rates are the lowest observed on-demand rows in data/providers.json, refreshed daily.

GPUVRAMFP16 fitINT4 fitLowest on-demand
H200 SXM5141 GBOOMOOM$2.79/hr (Vast.ai)
B200 Blackwell192 GBOOMOOM$3.99/hr (Vast.ai)
H100 SXM580 GBOOMOOM$1.89/hr (Vast.ai)
A100 80GB SXM480 GBOOMOOM$1.59/hr (Lambda Labs)
L40S48 GBOOMOOM$0.69/hr (Vast.ai)
RTX 409024 GBOOMOOM$0.34/hr (Vast.ai)

Recommended GPUs for DeepSeek V3

Minimum = smallest single GPU that fits the full INT4 stack (381.0 GB). Optimal = registry recommendation, falling back to the lowest observed on-demand rate among fitting cards. Every card links to its canonical GPU page.

No single GPU in the candidate set fits 381 GB INT4 โ€” DeepSeek V3 needs multi-GPU or a smaller model. Size it in the calculator โ†’

Workload Scenarios Using DeepSeek V3

Costed workload guides that price DeepSeek V3 against verified cloud rates and registry VRAM math.

Next steps

Frequently Asked Questions

How much VRAM does DeepSeek V3 need?โ–พ
DeepSeek V3 needs 1340 GB at FP16 (weights) and 336 GB at INT4 from the param-count weight model. Full-stack totals at 128K context are 1498.4 GB (FP16) and 381.0 GB (INT4) including KV-cache, CUDA overhead, and fragmentation headroom.
Can DeepSeek V3 run on a single GPU?โ–พ
Yes โ€” the registry recommendation is 8x H200, and the NVIDIA H200 SXM5 (141GB HBM3e) fits the FP16 workload single-GPU. FP16 full-stack needs 1498.4 GB.
What context length fits DeepSeek V3?โ–พ
Registry context window is 128,000 tokens. KV-cache grows linearly with context โ€” the quantization table shows totals at 4K, 33K, 128K context for FP16 and INT4.
Weight model: parameters ร— bytes-per-parameter. KV-cache: GQA formula or modeled estimate where layer geometry is unpublished.Methodology โ†’