Mixture of Experts (MoE) Explained: Memory and Serving Math
What Mixture of Experts (MoE) means for GPU memory: why 671B total but 37B active still needs all weights resident, plus bandwidth math.
Direct answer
A Mixture-of-Experts (MoE) model splits each transformer layer's feed-forward block into many expert sub-networks and routes each token through a small subset of them. The consequence for hardware is two separate budgets: resident capacity scales with the total parameter count (every expert's weights sit in VRAM whether or not it fires), while per-token weight reads scale with the active parameter count. DeepSeek R1 โ 671B total, 37B active per token, 5.5% โ still needs 671 GB of FP8 weights resident, but each token only streams about 37 GB of them: 7.7 ms of HBM time per token on an H200's 4.8 TB/s, against 14.7 ms for a dense 70B at FP8. The registry's MoE inventory, weights by arithmetic (run 2026-10-09):
| Model | Total params | Active/token | Active % | FP16 weights | FP8 weights | INT4 weights |
|---|---|---|---|---|---|---|
| DeepSeek R1 | 671B | 37B | 5.5% | 1,342 GB | 671 GB | 335.5 GB |
| DeepSeek V3 | 671B | 37B | 5.5% | 1,342 GB | 671 GB | 335.5 GB |
| Mixtral 8x22B | 141B | 22B | 15.6% | 282 GB | 141 GB | 70.5 GB |
| Nemotron-4 340B | 340B | 30B | 8.8% | 680 GB | 340 GB | 170 GB |
| Mistral Large | 123B | 12B | 9.8% | 246 GB | 123 GB | 61.5 GB |
| Ling 3.0 Flash | 124B | 5.1B | 4.1% | 248 GB | 124 GB | 62 GB |
| GLM-4 Flash | 10B | 10B | 100% | 20 GB | 10 GB | 5 GB |
| Llama 3.3 70B (dense, for scale) | 70.6B | 70.6B | 100% | 141.2 GB | 70.6 GB | 35.3 GB |
Source: OpenGPU Radar model registry (total/active parameters) ร bytes/param (FP16 2.0, FP8 1.0, INT4 0.5); run 2026-10-09. Weights only โ KV-cache, runtime and the 10% margin are added in the capacity table below.
How routing works โ and what "active" buys
Each MoE layer contains several expert networks of the same shape (a router scores the token and activates a fixed top-k of them; the rest sit idle for that token). The experts' weights are ordinary weights: they are loaded at startup and stay resident. This is the line the site's own DeepSeek pages keep returning to โ the minimum viable cluster article puts it as "37B of 671B parameters active per token โ but all expert weights must be resident," and the R1 cost article builds its whole economics case on it.
So MoE buys one thing and costs another:
- Buys: more total capacity per active FLOP โ the 671B model behaves, per token, like a ~37B one in weight-read terms.
- Costs: every byte of the 671B has to live somewhere โ capacity is billed on the total, compute and bandwidth on the active.
Two budgets, two different tables. Everything below separates them.
Capacity: you still pay for every parameter
In-service totals โ weights + KV-cache + 1.6 GB runtime + 10% fragmentation margin, batch 1, 4,096-token context โ from the VRAM canonical engine (run 2026-10-09):
| Model | FP16 total | FP8 total | INT4 total | Single-GPU read |
|---|---|---|---|---|
| DeepSeek R1 / V3 | 1,479.4 GB | 741.3 GB | 372.3 GB | none โ multi-GPU (registry: 8ร H200) |
| Mixtral 8x22B | 312.3 GB | 157.0 GB | 79.5 GB | INT4 marginal on 80 GB; FP8 fits 2ร80 GB |
| Nemotron-4 340B | 750.5 GB | 376.1 GB | 189.1 GB | none โ INT4 fits 2ร H200 (282 GB) |
| Mistral Large | 272.6 GB | 137.2 GB | 69.5 GB | INT4 fits 80 GB; FP8 marginal on H200 (141 GB) |
| Ling 3.0 Flash | 274.8 GB | 138.3 GB | 70.1 GB | INT4 fits 80 GB; FP8 marginal on H200 |
| GLM-4 Flash | 23.8 GB | 12.8 GB | 7.3 GB | INT4 fits 8 GB and up |
| Llama 3.3 70B (dense) | 157.6 GB | 79.7 GB | 40.8 GB | INT4 fits 48 GB; FP8 marginal on 80 GB |
Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-09; batch 1, single GPU, 4,096-token context. Bold = fits one 80 GB GPU (Marginal when headroom < 10%).
Read it the way the VRAM tiers article reads its tables โ with capacity, not activity, as the constraint: Mixtral's 22B active tokens never change the fact that its INT4 build lands at 79.5 GB, a 0.5 GB-margin fit on an 80 GB card. Context is the multiplier after that: R1 at the full 128K window adds its compressed MLA KV-cache (41.9 GB, medium-confidence engine estimate) for 786 GB FP8 / 417 GB INT4 in service โ the same figures the R1 cost article publishes and the reason the registry recommends 8ร H200. KV mechanics: KV-cache explained.
Bandwidth: what each token actually streams
Decode is memory-bound: each new token re-reads weights from HBM. Weight-read time = active weight GB รท bandwidth โ an arithmetic floor at batch 1 (lower bound on per-token latency; real serving adds KV reads, attention and scheduling and lands slower):
| Weights read per token | On H200 (4.8 TB/s) | On H100 (3.35 TB/s) |
|---|---|---|
| R1 active weights, FP8 (37 GB) | 7.7 ms | 11.0 ms |
| Dense 70B weights, FP8 (70.6 GB) | 14.7 ms | 21.1 ms |
| R1 active weights, INT4 (18.5 GB) | 3.9 ms | 5.5 ms |
Source: arithmetic โ registry active-parameter count ร bytes/param รท manufacturer bandwidth (H200 4.8 TB/s, H100 3.35 TB/s per GPU specs); run 2026-10-09. CALCULATED_ESTIMATE: weights-streaming floor only, batch 1.
Two readings. First, on the same GPU at the same precision, R1 pays ~1.9ร less weight-read time per token than a dense 70B (7.7 vs 14.7 ms on H200) โ that gap is MoE's per-token bandwidth advantage in one number. Second, halving weights with precision halves the read: INT4 active reads (18.5 GB) run 3.9 ms on H200, so quantization does double duty for MoE โ it shrinks the resident-capacity bill and the per-token stream, at the INT4 quality trade-off the quantization explainer documents.
No throughput claims follow from this floor, deliberately: the cost article's stance is that multi-GPU R1 throughput is PROJECTED territory unless it comes from a real cluster record, and nothing here changes that.
What changes at batch > 1
The two-budget model is batch-dependent, and it's worth being explicit:
- Dense: each forward step reads the full weight set once regardless of batch size โ more tokens per step amortize the same read.
- MoE at batch 1: reads โ active weights per token (the floor above).
- MoE at large batch: with enough concurrent tokens, nearly every expert is hit in each step โ per-step reads approach the total weight set, and MoE's per-token bandwidth advantage shrinks toward the dense amortization curve. The resident-capacity bill never changes: all experts were resident at step one.
This is why the site's DeepSeek pages treat MoE decode as memory-bandwidth bound (cost article) and why capacity โ not activity โ drives cluster sizing (minimum viable cluster). The one throughput figure the site publishes (100 tok/s, from its own cluster record) belongs to that page, not to this arithmetic.
FAQ
If only 37B parameters are active, why can't I run DeepSeek R1 on a 2รH200 card's worth of memory? Because routing is per token, not per lifetime: over a conversation, every expert fires, and an expert that isn't resident can't fire at all. INT4 gets the resident bill to 372.3 GB at 4K โ still multi-GPU โ and to 417 GB at 128K.
Is a bigger active ratio worse? It moves the model toward dense behavior: more weight-read per token (the bandwidth table) but fewer total weights for the same active compute. The ratio is a design trade, not a quality dial โ nothing in this article's arithmetic says which routing choice is better.
Does MoE change anything about quantization? No โ bytes/param is bytes/param. FP8 halves, INT4 quarters both the resident capacity (capacity table) and the per-token read (bandwidth table). The quality trade-offs are the quantization guide's to make: FP16 vs FP8 vs INT4.
Which of these can I actually run locally? Per the capacity table: GLM-4 Flash INT4 fits an 8 GB card; Mistral Large and Ling 3.0 Flash INT4 fit 80 GB; Mixtral 8x22B INT4 is a marginal 80 GB fit (79.5 GB). Full sizing across context lengths: VRAM tiers and the calculator.
Limitations and assumptions
- Registry figures, arithmetic propagation. Total/active parameters are the OpenGPU Radar model registry's values as published on each model page; weights = parameters ร bytes/param; percentages = active รท total. Where a registry active figure looks surprising (GLM-4 Flash lists active = total), we publish the registry's numbers rather than our own guesses.
- Capacity table conditions: batch 1, single GPU, 4,096-token context, run 2026-10-09; totals = weights + KV + 1.2 GB runtime + 0.4 GB activation + 10% margin โ identical methodology to the VRAM tiers article. KV confidence: exact GQA (high) for the dense anchor; MLA modeled estimate (medium) for DeepSeek; estimated models (medium) for Mixtral, Nemotron, Mistral Large, Ling and GLM โ engine output, only the confidence differs.
- Bandwidth table: weights-streaming floor only โ excludes KV reads, attention, all-reduce, scheduler overhead, and speculative decoding; manufacturer bandwidth specs; batch 1. No tok/s, no benchmarks, no prices.
- Not covered: training MoE models, expert parallelism, router loss/load-balancing mechanics, and provider-level pricing (rates live on the GPU pages).
Related resources
- VRAM Calculator ยท Which LLMs fit 8โ80GB GPUs ยท Model hub
- DeepSeek: R1 cost breakdown ยท Minimum viable cluster ยท R1 FP8 vs INT4 memory ยท Calculator requirements
- Concepts: KV-cache explained ยท Quantization explained ยท Local serving engines
- GPUs: H200 ยท H100 ยท RTX 4090 ยท L40S ยท B200
Related Articles
B300 Blackwell Ultra: Full Architecture Deep Dive and VRAM Analysis
B300 Blackwell Ultra: 288GB HBM3e, 9 TB/s, 2,500 FP8 TFLOPS โ NVIDIA's target spec per the Blackwell Ultra whitepaper. VRAM analysis: what fits on 3x B300 vs 5x B200, and why pre-production means verify before buying.
ArchitectureDeepSeek R1 671B: FP8 vs INT4 Memory Overhead & PPL Degradation
DeepSeek R1 671B: 786 GB VRAM (FP8, 128K ctx) vs 417 GB (INT4) on H200/B200. Memory overhead breakdown, precision tradeoffs, and observed cluster pricing.