VRAM Reference

Phi-4 14B VRAM Requirements: FP16 / INT4 / KV-Cache (2026)

Phi-4 14B runs on 14B parameters with a 16K-token context window. Weights are deterministic param math; KV-cache uses the registry layer geometry where published, otherwise a modeled GQA estimate โ€” every figure is provenance-labeled below.

FP16 weights28 GB
INT4 weights10 GB
Context16K tokens
Single-GPU verdictFits (INT4, 10 GB)
28 GB FP16 weights
Calculated EstimateHIGH
SourceOpenGPU Radar models-registry.json (param-count weight model)
VerifiedSep 29, 2026
Value
28 GB
Methodology

weights = parametersB ร— bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)

Refreshed daily from provider APIs and market scrapingSep 29, 2026
10 GB INT4 weights
Calculated EstimateHIGH
SourceOpenGPU Radar models-registry.json (param-count weight model)
VerifiedSep 29, 2026
Value
10 GB
Methodology

weights = parametersB ร— bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)

Refreshed daily from provider APIs and market scrapingSep 29, 2026
16K context window
Manufacturer SpecHIGH
SourceModel vendor context-window specification
VerifiedSep 29, 2026
Value
16,000 tokens
Refreshed daily from provider APIs and market scrapingSep 29, 2026
Methodology โ†’

How much VRAM does Phi-4 14B need?

Weights: 28 GB at FP16, 10 GB at INT4 (parameters ร— bytes-per-parameter: FP16 = 2 B, INT4 = 0.5 B).

Full stack at 16K: 32.7 GB FP16 / 9.5 GB INT4 โ€” including KV-cache, CUDA overhead, activations, and fragmentation headroom from the canonical VRAM engine.

Single-GPU verdict: INT4 fits in 24 GB class cards (e.g. RTX 4090); FP16 needs 80 GB-class hardware.

VRAM by Precision & Context

From the canonical VRAM engine: weights + KV-cache + CUDA overhead + activations + fragmentation headroom. KV-cache uses published layer geometry where available; otherwise a modeled GQA estimate (labeled in the badge).

PrecisionContextWeightsKV-CacheTotal (est.)
FP164K28.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
28 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 14
  • quantization: fp16
  • bytesPerParam: 2
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.03 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.03 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 4096
  • quantization: fp16
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
32.6 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
32.59 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 14
  • quantization: fp16
  • contextLength: 4096
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: phi-4-14b
  • bytesPerParam: 2
  • bytesPerKvElement: 2
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026
FP1616K28.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
28 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 14
  • quantization: fp16
  • bytesPerParam: 2
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.10 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.1 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 16000
  • quantization: fp16
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
32.7 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
32.67 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 14
  • quantization: fp16
  • contextLength: 16000
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: phi-4-14b
  • bytesPerParam: 2
  • bytesPerKvElement: 2
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026
INT44K7.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
7 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 14
  • quantization: int4
  • bytesPerParam: 0.5
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.01 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.01 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 4096
  • quantization: int4
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
9.5 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
9.47 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 14
  • quantization: int4
  • contextLength: 4096
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: phi-4-14b
  • bytesPerParam: 0.5
  • bytesPerKvElement: 1
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026
INT416K7.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
7 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 14
  • quantization: int4
  • bytesPerParam: 0.5
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.05 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.05 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 16000
  • quantization: int4
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
9.5 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
9.52 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 14
  • quantization: int4
  • contextLength: 16000
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: phi-4-14b
  • bytesPerParam: 0.5
  • bytesPerKvElement: 1
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026

Cloud GPUs That Fit Phi-4 14B

Fit = full-stack total (weights + KV at 16K) โ‰ค GPU VRAM. Rates are the lowest observed on-demand rows in data/providers.json, refreshed daily.

GPUVRAMFP16 fitINT4 fitLowest on-demand
H200 SXM5141 GBโœ“ fitsโœ“ fits$2.79/hr (Vast.ai)
B200 Blackwell192 GBโœ“ fitsโœ“ fits$3.99/hr (Vast.ai)
H100 SXM580 GBโœ“ fitsโœ“ fits$1.89/hr (Vast.ai)
A100 80GB SXM480 GBโœ“ fitsโœ“ fits$1.59/hr (Lambda Labs)
L40S48 GBโœ“ fitsโœ“ fits$0.69/hr (Vast.ai)
RTX 409024 GBOOMโœ“ fits$0.34/hr (Vast.ai)

Recommended GPUs for Phi-4 14B

Minimum = smallest single GPU that fits the full INT4 stack (9.5 GB). Optimal = registry recommendation, falling back to the lowest observed on-demand rate among fitting cards. Every card links to its canonical GPU page.

Next steps

Frequently Asked Questions

How much VRAM does Phi-4 14B need?โ–พ
Phi-4 14B needs 28 GB at FP16 (weights) and 10 GB at INT4 from the param-count weight model. Full-stack totals at 16K context are 32.7 GB (FP16) and 9.5 GB (INT4) including KV-cache, CUDA overhead, and fragmentation headroom.
Can Phi-4 14B run on a single GPU?โ–พ
Yes โ€” the registry recommendation is 1x L40S, and the NVIDIA L40S (48GB GDDR6) fits the INT4 workload single-GPU. FP16 full-stack needs 32.7 GB.
What context length fits Phi-4 14B?โ–พ
Registry context window is 16,000 tokens. KV-cache grows linearly with context โ€” the quantization table shows totals at 4K, 16K context for FP16 and INT4.
Weight model: parameters ร— bytes-per-parameter. KV-cache: GQA formula or modeled estimate where layer geometry is unpublished.Methodology โ†’