VRAM Reference

Gemma 2 27B VRAM Requirements: FP16 / INT4 / KV-Cache (2026)

Gemma 2 27B runs on 27B parameters with a 8K-token context window. Weights are deterministic param math; KV-cache uses the registry layer geometry where published, otherwise a modeled GQA estimate โ€” every figure is provenance-labeled below.

FP16 weights54 GB
INT4 weights16 GB
Context8K tokens
Single-GPU verdictFits (INT4, 17 GB)
54 GB FP16 weights
Calculated EstimateHIGH
SourceOpenGPU Radar models-registry.json (param-count weight model)
VerifiedSep 29, 2026
Value
54 GB
Methodology

weights = parametersB ร— bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)

Refreshed daily from provider APIs and market scrapingSep 29, 2026
16 GB INT4 weights
Calculated EstimateHIGH
SourceOpenGPU Radar models-registry.json (param-count weight model)
VerifiedSep 29, 2026
Value
16 GB
Methodology

weights = parametersB ร— bytes-per-parameter (FP16 = 2 B, INT4 = 0.5 B)

Refreshed daily from provider APIs and market scrapingSep 29, 2026
8K context window
Manufacturer SpecHIGH
SourceModel vendor context-window specification
VerifiedSep 29, 2026
Value
8,000 tokens
Refreshed daily from provider APIs and market scrapingSep 29, 2026
Methodology โ†’

How much VRAM does Gemma 2 27B need?

Weights: 54 GB at FP16, 16 GB at INT4 (parameters ร— bytes-per-parameter: FP16 = 2 B, INT4 = 0.5 B).

Full stack at 8K: 61.3 GB FP16 / 16.7 GB INT4 โ€” including KV-cache, CUDA overhead, activations, and fragmentation headroom from the canonical VRAM engine.

Single-GPU verdict: INT4 fits in 24 GB class cards (e.g. RTX 4090); FP16 needs 80 GB-class hardware.

VRAM by Precision & Context

From the canonical VRAM engine: weights + KV-cache + CUDA overhead + activations + fragmentation headroom. KV-cache uses published layer geometry where available; otherwise a modeled GQA estimate (labeled in the badge).

PrecisionContextWeightsKV-CacheTotal (est.)
FP164K54.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
54 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 27
  • quantization: fp16
  • bytesPerParam: 2
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.05 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.05 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 4096
  • quantization: fp16
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
61.2 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
61.22 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 27
  • quantization: fp16
  • contextLength: 4096
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: gemma-2-27b
  • bytesPerParam: 2
  • bytesPerKvElement: 2
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026
FP168K54.0 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
54 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: FP16 = 2 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 27
  • quantization: fp16
  • bytesPerParam: 2
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.10 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.1 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 8000
  • quantization: fp16
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
61.3 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
61.27 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 27
  • quantization: fp16
  • contextLength: 8000
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: gemma-2-27b
  • bytesPerParam: 2
  • bytesPerKvElement: 2
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026
INT44K13.5 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
13.5 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 27
  • quantization: int4
  • bytesPerParam: 0.5
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.03 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.03 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 4096
  • quantization: int4
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
16.6 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
16.64 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 27
  • quantization: int4
  • contextLength: 4096
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: gemma-2-27b
  • bytesPerParam: 0.5
  • bytesPerKvElement: 1
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026
INT48K13.5 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
13.5 GB
Methodology

Formula: parameterCountB ร— bytesPerParam. Precision: INT4 = 0.5 bytes/param. (parameterCountB in billions, result in GB).

Assumptions & Parameters
  • paramCountB: 27
  • quantization: int4
  • bytesPerParam: 0.5
Refreshed daily from provider APIs and market scrapingSep 26, 2026
0.05 GB
Calculated EstimateMEDIUM
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
0.05 GB
Methodology

Estimated via GQA model โ€” architectural parameters not available.

Assumptions & Parameters
  • contextLength: 8000
  • quantization: int4
  • estimatedVia: Modeled GQA estimate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
16.7 GB
Calculated EstimateHIGH
SourceDeterministic VRAM Canonical Engine
VerifiedSep 26, 2026
Value
16.66 GB
Methodology

Total = weights + kvCache + cudaOverhead + activation + fragmentationHeadroom (10%).

Assumptions & Parameters
  • paramCountB: 27
  • quantization: int4
  • contextLength: 8000
  • batchSize: 1
  • numLayers: 0
  • numKvHeads: 0
  • headDim: 0
  • gqaRatio: 1
  • architecture: gemma-2-27b
  • bytesPerParam: 0.5
  • bytesPerKvElement: 1
  • cudaOverheadGb: 1.2
  • activationGb: 0.4
  • fragmentationHeadroomPct: 10
  • isMla: false
Refreshed daily from provider APIs and market scrapingSep 26, 2026

Cloud GPUs That Fit Gemma 2 27B

Fit = full-stack total (weights + KV at 8K) โ‰ค GPU VRAM. Rates are the lowest observed on-demand rows in data/providers.json, refreshed daily.

GPUVRAMFP16 fitINT4 fitLowest on-demand
H200 SXM5141 GBโœ“ fitsโœ“ fits$2.79/hr (Vast.ai)
B200 Blackwell192 GBโœ“ fitsโœ“ fits$3.99/hr (Vast.ai)
H100 SXM580 GBโœ“ fitsโœ“ fits$1.89/hr (Vast.ai)
A100 80GB SXM480 GBโœ“ fitsโœ“ fits$1.59/hr (Lambda Labs)
L40S48 GBOOMโœ“ fits$0.69/hr (Vast.ai)
RTX 409024 GBOOMโœ“ fits$0.34/hr (Vast.ai)

Recommended GPUs for Gemma 2 27B

Minimum = smallest single GPU that fits the full INT4 stack (16.7 GB). Optimal = registry recommendation, falling back to the lowest observed on-demand rate among fitting cards. Every card links to its canonical GPU page.

Next steps

Frequently Asked Questions

How much VRAM does Gemma 2 27B need?โ–พ
Gemma 2 27B needs 54 GB at FP16 (weights) and 16 GB at INT4 from the param-count weight model. Full-stack totals at 8K context are 61.3 GB (FP16) and 16.7 GB (INT4) including KV-cache, CUDA overhead, and fragmentation headroom.
Can Gemma 2 27B run on a single GPU?โ–พ
Yes โ€” the registry recommendation is 1x H200, and the NVIDIA H200 SXM5 (141GB HBM3e) fits the INT4 workload single-GPU. FP16 full-stack needs 61.3 GB.
What context length fits Gemma 2 27B?โ–พ
Registry context window is 8,000 tokens. KV-cache grows linearly with context โ€” the quantization table shows totals at 4K, 8K context for FP16 and INT4.
Weight model: parameters ร— bytes-per-parameter. KV-cache: GQA formula or modeled estimate where layer geometry is unpublished.Methodology โ†’