Mistral Large VRAM Requirements: FP16 / INT4 / KV-Cache (2026)
Mistral Large runs on 123B parameters with a 128K-token context window. Weights are deterministic param math; KV-cache uses the registry layer geometry where published, otherwise a modeled GQA estimate โ every figure is provenance-labeled below.
74 GB FP16 weights
weights = parametersB ร bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)
37 GB INT4 weights
weights = parametersB ร bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)
128K context window
How much VRAM does Mistral Large need?
Weights: 74 GB at FP16, 37 GB at INT4 (parameters ร bytes-per-parameter: FP16 = 2 B, INT4 = 0.5 B).
Full stack at 128K: 280.3 GB FP16 / 73.4 GB INT4 โ including KV-cache, CUDA overhead, activations, and fragmentation headroom from the canonical VRAM engine.
Single-GPU verdict: INT4 fits 80 GB datacenter cards (A100/H100); FP16 needs 320 GB โ multi-GPU.
VRAM by Precision & Context
From the canonical VRAM engine: weights + KV-cache + CUDA overhead + activations + fragmentation headroom. KV-cache uses published layer geometry where available; otherwise a modeled GQA estimate (labeled in the badge).
| Precision | Context | Weights | KV-Cache | Total (est.) |
|---|---|---|---|---|
| FP16 | 4K | 246.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 246 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.23 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.23 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 272.6 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 272.62 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| FP16 | 33K | 246.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 246 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 1.85 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 1.85 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 274.4 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 274.4 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| FP16 | 128K | 246.0 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 246 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 7.24 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 7.24 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 280.3 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 280.33 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| INT4 | 4K | 61.5 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 61.5 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.12 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.12 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 69.5 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 69.54 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| INT4 | 33K | 61.5 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 61.5 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 0.93 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 0.93 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 70.4 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 70.43 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
| INT4 | 128K | 61.5 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 61.5 GB Methodology Formula: parameterCountB ร bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 3.62 GB Calculated EstimateMEDIUM SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 3.62 GB Methodology Estimated via GQA model โ architectural parameters not available. Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 | 73.4 GB Calculated EstimateHIGH SourceDeterministic VRAM Canonical Engine VerifiedSep 26, 2026 Value 73.39 GB Methodology Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%). Assumptions & Parameters
Refreshed daily from provider APIs and market scrapingSep 26, 2026 |
Cloud GPUs That Fit Mistral Large
Fit = full-stack total (weights + KV at 128K) โค GPU VRAM. Rates are the lowest observed on-demand rows in data/providers.json, refreshed daily.
| GPU | VRAM | FP16 fit | INT4 fit | Lowest on-demand |
|---|---|---|---|---|
| H200 SXM5 | 141 GB | OOM | โ fits | $2.79/hr (Vast.ai) |
| B200 Blackwell | 192 GB | OOM | โ fits | $3.99/hr (Vast.ai) |
| H100 SXM5 | 80 GB | OOM | โ fits | $1.89/hr (Vast.ai) |
| A100 80GB SXM4 | 80 GB | OOM | โ fits | $1.59/hr (Lambda Labs) |
| L40S | 48 GB | OOM | OOM | $0.69/hr (Vast.ai) |
| RTX 4090 | 24 GB | OOM | OOM | $0.34/hr (Vast.ai) |
Recommended GPUs for Mistral Large
Minimum = smallest single GPU that fits the full INT4 stack (73.4 GB). Optimal = registry recommendation, falling back to the lowest observed on-demand rate among fitting cards. Every card links to its canonical GPU page.
Next steps