Phi-4 14B VRAM Requirements: FP16 / INT4 / KV-Cache (2026)
Phi-4 14B runs on 14B parameters with a 16K-token context window. Weights are deterministic param math; KV-cache uses the registry layer geometry where published, otherwise a modeled GQA estimate โ every figure is provenance-labeled below.
28 GB FP16 weights
weights = parametersB ร bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)
10 GB INT4 weights
weights = parametersB ร bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)
16K context window
How much VRAM does Phi-4 14B need?
Weights: 28 GB at FP16, 10 GB at INT4 (parameters ร bytes-per-parameter: FP16 = 2 B, INT4 = 0.5 B).
Full stack at 16K: 32.7 GB FP16 / 9.5 GB INT4 โ including KV-cache, CUDA overhead, activations, and fragmentation headroom from the canonical VRAM engine.
Single-GPU verdict: INT4 fits in 24 GB class cards (e.g. RTX 4090); FP16 needs 80 GB-class hardware.
VRAM by Precision & Context
From the canonical VRAM engine: weights + KV-cache + CUDA overhead + activations + fragmentation headroom. KV-cache uses published layer geometry where available; otherwise a modeled GQA estimate (labeled in the badge).
| Precision | Context | Weights | KV-Cache | Total (est.) |
|---|---|---|---|---|
| FP16 | 4K | 28.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 28 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.03 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.03 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 32.6 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 32.59 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| FP16 | 16K | 28.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 28 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.10 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.1 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 32.7 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 32.67 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| INT4 | 4K | 7.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 7 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.01 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.01 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 9.5 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 9.47 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| INT4 | 16K | 7.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 7 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.05 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.05 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 9.5 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 9.52 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
Cloud GPUs That Fit Phi-4 14B
Fit = full-stack total (weights + KV at 16K) โค GPU VRAM. Rates are the lowest observed on-demand rows in data/providers.json, refreshed daily.
| GPU | VRAM | FP16 fit | INT4 fit | Lowest on-demand |
|---|---|---|---|---|
| H200 SXM5 | 141 GB | โ fits | โ fits | $2.79/hr (Vast.ai) |
| B200 Blackwell | 192 GB | โ fits | โ fits | $3.99/hr (Vast.ai) |
| H100 SXM5 | 80 GB | โ fits | โ fits | $1.89/hr (Vast.ai) |
| A100 80GB SXM4 | 80 GB | โ fits | โ fits | $1.59/hr (Lambda Labs) |
| L40S | 48 GB | โ fits | โ fits | $0.69/hr (Vast.ai) |
| RTX 4090 | 24 GB | OOM | โ fits | $0.34/hr (Vast.ai) |
Recommended GPUs for Phi-4 14B
Minimum = smallest single GPU that fits the full INT4 stack (9.5 GB). Optimal = registry recommendation, falling back to the lowest observed on-demand rate among fitting cards. Every card links to its canonical GPU page.
Next steps