VRAM Reference

Codestral 22B VRAM Requirements: FP16 / INT4 / KV-Cache (2026)

Codestral 22B runs on 22B parameters with a 32K-token context window. Weights are deterministic param math; KV-cache uses the registry layer geometry where published, otherwise a modeled GQA estimate โ€” every figure is provenance-labeled below.

FP16 weights44 GB
INT4 weights14 GB
Context32K tokens
Single-GPU verdictFits (INT4, 14 GB)
44 GB FP16 weights
Calculated EstimateHIGH
SourceOpenGPU Radar models-registry.json (param-count weight model)
VerifiedSep 29, 2026
Value
44 GB
Methodology

weights = parametersB ร— bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)

Refreshed daily from provider APIs and market scrapingSep 29, 2026
14 GB INT4 weights
Calculated EstimateHIGH
SourceOpenGPU Radar models-registry.json (param-count weight model)
VerifiedSep 29, 2026
Value
14 GB
Methodology

weights = parametersB ร— bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)

Refreshed daily from provider APIs and market scrapingSep 29, 2026
32K context window
Manufacturer SpecHIGH
SourceModel vendor context-window specification
VerifiedSep 29, 2026
Value
32,000 tokens
Refreshed daily from provider APIs and market scrapingSep 29, 2026
Methodology โ†’

How much VRAM does Codestral 22B need?

Weights: 44 GB at FP16, 14 GB at INT4 (parameters ร— bytes-per-parameter: FP16 = 2 B, INT4 = 0.5 B).

Full stack at 32K: 50.5 GB FP16 / 14.0 GB INT4 โ€” including KV-cache, CUDA overhead, activations, and fragmentation headroom from the canonical VRAM engine.

Single-GPU verdict: INT4 fits in 24 GB class cards (e.g. RTX 4090); FP16 needs 80 GB-class hardware.

VRAM by Precision & Context

From the canonical VRAM engine: weights + KV-cache + CUDA overhead + activations + fragmentation headroom. KV-cache uses published layer geometry where available; otherwise a modeled GQA estimate (labeled in the badge).

PrecisionContextWeightsKV-CacheTotal (est.)
FP164K44.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
44 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 22
  • quantization: fp16
  • bytesPerParam: 2
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.04 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.04 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 4096
  • quantization: fp16
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
50.2 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
50.2 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 22
  • quantization: fp16
  • contextLength: 4096
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: codestral-22b
  • bytesPerParam: 2
  • bytesPerKvElement: 2
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026
FP1632K44.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
44 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 22
  • quantization: fp16
  • bytesPerParam: 2
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.32 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.32 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 32000
  • quantization: fp16
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
50.5 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
50.52 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 22
  • quantization: fp16
  • contextLength: 32000
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: codestral-22b
  • bytesPerParam: 2
  • bytesPerKvElement: 2
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026
INT44K11.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
11 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 22
  • quantization: int4
  • bytesPerParam: 0.5
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.02 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.02 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 4096
  • quantization: int4
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
13.9 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
13.88 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 22
  • quantization: int4
  • contextLength: 4096
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: codestral-22b
  • bytesPerParam: 0.5
  • bytesPerKvElement: 1
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026
INT432K11.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
11 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 22
  • quantization: int4
  • bytesPerParam: 0.5
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.16 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.16 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 32000
  • quantization: int4
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
14.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
14.04 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 22
  • quantization: int4
  • contextLength: 32000
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: codestral-22b
  • bytesPerParam: 0.5
  • bytesPerKvElement: 1
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026

Cloud GPUs That Fit Codestral 22B

Fit = full-stack total (weights + KV at 32K) โ‰ค GPU VRAM. Rates are the lowest observed on-demand rows in data/providers.json, refreshed daily.

GPUVRAMFP16 fitINT4 fitLowest on-demand
H200 SXM5141 GBโœ“ fitsโœ“ fits$2.79/hr (Vast.ai)
B200 Blackwell192 GBโœ“ fitsโœ“ fits$3.99/hr (Vast.ai)
H100 SXM580 GBโœ“ fitsโœ“ fits$1.89/hr (Vast.ai)
A100 80GB SXM480 GBโœ“ fitsโœ“ fits$1.59/hr (Lambda Labs)
L40S48 GBOOMโœ“ fits$0.69/hr (Vast.ai)
RTX 409024 GBOOMโœ“ fits$0.34/hr (Vast.ai)

Recommended GPUs for Codestral 22B

Minimum = smallest single GPU that fits the full INT4 stack (14.0 GB). Optimal = registry recommendation, falling back to the lowest observed on-demand rate among fitting cards. Every card links to its canonical GPU page.

Next steps

Frequently Asked Questions

How much VRAM does Codestral 22B need?โ–พ
Codestral 22B needs 44 GB at FP16 (weights) and 14 GB at INT4 from the param-count weight model. Full-stack totals at 32K context are 50.5 GB (FP16) and 14.0 GB (INT4) including KV-cache, CUDA overhead, and fragmentation headroom.
Can Codestral 22B run on a single GPU?โ–พ
Yes โ€” the registry recommendation is 1x H200, and the NVIDIA H200 SXM5 (141GB HBM3e) fits the INT4 workload single-GPU. FP16 full-stack needs 50.5 GB.
What context length fits Codestral 22B?โ–พ
Registry context window is 32,000 tokens. KV-cache grows linearly with context โ€” the quantization table shows totals at 4K, 32K context for FP16 and INT4.
Weight model: parameters ร— bytes-per-parameter. KV-cache: GQA formula or modeled estimate where layer geometry is unpublished.Methodology โ†’