Codestral 22B VRAM Requirements: FP16 / INT4 / KV-Cache (2026)
Codestral 22B runs on 22B parameters with a 32K-token context window. Weights are deterministic param math; KV-cache uses the registry layer geometry where published, otherwise a modeled GQA estimate โ every figure is provenance-labeled below.
44 GB FP16 weights
weights = parametersB ร bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)
14 GB INT4 weights
weights = parametersB ร bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)
32K context window
How much VRAM does Codestral 22B need?
Weights: 44 GB at FP16, 14 GB at INT4 (parameters ร bytes-per-parameter: FP16 = 2 B, INT4 = 0.5 B).
Full stack at 32K: 50.5 GB FP16 / 14.0 GB INT4 โ including KV-cache, CUDA overhead, activations, and fragmentation headroom from the canonical VRAM engine.
Single-GPU verdict: INT4 fits in 24 GB class cards (e.g. RTX 4090); FP16 needs 80 GB-class hardware.
VRAM by Precision & Context
From the canonical VRAM engine: weights + KV-cache + CUDA overhead + activations + fragmentation headroom. KV-cache uses published layer geometry where available; otherwise a modeled GQA estimate (labeled in the badge).
| Precision | Context | Weights | KV-Cache | Total (est.) |
|---|---|---|---|---|
| FP16 | 4K | 44.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 44 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.04 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.04 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 50.2 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 50.2 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| FP16 | 32K | 44.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 44 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.32 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.32 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 50.5 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 50.52 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| INT4 | 4K | 11.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 11 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.02 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.02 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 13.9 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 13.88 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| INT4 | 32K | 11.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 11 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.16 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.16 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 14.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 14.04 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
Cloud GPUs That Fit Codestral 22B
Fit = full-stack total (weights + KV at 32K) โค GPU VRAM. Rates are the lowest observed on-demand rows in data/providers.json, refreshed daily.
| GPU | VRAM | FP16 fit | INT4 fit | Lowest on-demand |
|---|---|---|---|---|
| H200 SXM5 | 141 GB | โ fits | โ fits | $2.79/hr (Vast.ai) |
| B200 Blackwell | 192 GB | โ fits | โ fits | $3.99/hr (Vast.ai) |
| H100 SXM5 | 80 GB | โ fits | โ fits | $1.89/hr (Vast.ai) |
| A100 80GB SXM4 | 80 GB | โ fits | โ fits | $1.59/hr (Lambda Labs) |
| L40S | 48 GB | OOM | โ fits | $0.69/hr (Vast.ai) |
| RTX 4090 | 24 GB | OOM | โ fits | $0.34/hr (Vast.ai) |
Recommended GPUs for Codestral 22B
Minimum = smallest single GPU that fits the full INT4 stack (14.0 GB). Optimal = registry recommendation, falling back to the lowest observed on-demand rate among fitting cards. Every card links to its canonical GPU page.
Next steps