Architecture2026-10-09โ€ขBy Sreeโ€ข5 min read

Mixture of Experts (MoE) Explained: Memory and Serving Math

What Mixture of Experts (MoE) means for GPU memory: why 671B total but 37B active still needs all weights resident, plus bandwidth math.

Direct answer

A Mixture-of-Experts (MoE) model splits each transformer layer's feed-forward block into many expert sub-networks and routes each token through a small subset of them. The consequence for hardware is two separate budgets: resident capacity scales with the total parameter count (every expert's weights sit in VRAM whether or not it fires), while per-token weight reads scale with the active parameter count. DeepSeek R1 โ€” 671B total, 37B active per token, 5.5% โ€” still needs 671 GB of FP8 weights resident, but each token only streams about 37 GB of them: 7.7 ms of HBM time per token on an H200's 4.8 TB/s, against 14.7 ms for a dense 70B at FP8. The registry's MoE inventory, weights by arithmetic (run 2026-10-09):

ModelTotal paramsActive/tokenActive %FP16 weightsFP8 weightsINT4 weights
DeepSeek R1671B37B5.5%1,342 GB671 GB335.5 GB
DeepSeek V3671B37B5.5%1,342 GB671 GB335.5 GB
Mixtral 8x22B141B22B15.6%282 GB141 GB70.5 GB
Nemotron-4 340B340B30B8.8%680 GB340 GB170 GB
Mistral Large123B12B9.8%246 GB123 GB61.5 GB
Ling 3.0 Flash124B5.1B4.1%248 GB124 GB62 GB
GLM-4 Flash10B10B100%20 GB10 GB5 GB
Llama 3.3 70B (dense, for scale)70.6B70.6B100%141.2 GB70.6 GB35.3 GB

Source: OpenGPU Radar model registry (total/active parameters) ร— bytes/param (FP16 2.0, FP8 1.0, INT4 0.5); run 2026-10-09. Weights only โ€” KV-cache, runtime and the 10% margin are added in the capacity table below.

How routing works โ€” and what "active" buys

Each MoE layer contains several expert networks of the same shape (a router scores the token and activates a fixed top-k of them; the rest sit idle for that token). The experts' weights are ordinary weights: they are loaded at startup and stay resident. This is the line the site's own DeepSeek pages keep returning to โ€” the minimum viable cluster article puts it as "37B of 671B parameters active per token โ€” but all expert weights must be resident," and the R1 cost article builds its whole economics case on it.

So MoE buys one thing and costs another:

  • Buys: more total capacity per active FLOP โ€” the 671B model behaves, per token, like a ~37B one in weight-read terms.
  • Costs: every byte of the 671B has to live somewhere โ€” capacity is billed on the total, compute and bandwidth on the active.

Two budgets, two different tables. Everything below separates them.

Capacity: you still pay for every parameter

In-service totals โ€” weights + KV-cache + 1.6 GB runtime + 10% fragmentation margin, batch 1, 4,096-token context โ€” from the VRAM canonical engine (run 2026-10-09):

ModelFP16 totalFP8 totalINT4 totalSingle-GPU read
DeepSeek R1 / V31,479.4 GB741.3 GB372.3 GBnone โ€” multi-GPU (registry: 8ร— H200)
Mixtral 8x22B312.3 GB157.0 GB79.5 GBINT4 marginal on 80 GB; FP8 fits 2ร—80 GB
Nemotron-4 340B750.5 GB376.1 GB189.1 GBnone โ€” INT4 fits 2ร— H200 (282 GB)
Mistral Large272.6 GB137.2 GB69.5 GBINT4 fits 80 GB; FP8 marginal on H200 (141 GB)
Ling 3.0 Flash274.8 GB138.3 GB70.1 GBINT4 fits 80 GB; FP8 marginal on H200
GLM-4 Flash23.8 GB12.8 GB7.3 GBINT4 fits 8 GB and up
Llama 3.3 70B (dense)157.6 GB79.7 GB40.8 GBINT4 fits 48 GB; FP8 marginal on 80 GB

Source: OpenGPU Radar VRAM canonical engine (calculateCanonicalVram), run 2026-10-09; batch 1, single GPU, 4,096-token context. Bold = fits one 80 GB GPU (Marginal when headroom < 10%).

Read it the way the VRAM tiers article reads its tables โ€” with capacity, not activity, as the constraint: Mixtral's 22B active tokens never change the fact that its INT4 build lands at 79.5 GB, a 0.5 GB-margin fit on an 80 GB card. Context is the multiplier after that: R1 at the full 128K window adds its compressed MLA KV-cache (41.9 GB, medium-confidence engine estimate) for 786 GB FP8 / 417 GB INT4 in service โ€” the same figures the R1 cost article publishes and the reason the registry recommends 8ร— H200. KV mechanics: KV-cache explained.

Bandwidth: what each token actually streams

Decode is memory-bound: each new token re-reads weights from HBM. Weight-read time = active weight GB รท bandwidth โ€” an arithmetic floor at batch 1 (lower bound on per-token latency; real serving adds KV reads, attention and scheduling and lands slower):

Weights read per tokenOn H200 (4.8 TB/s)On H100 (3.35 TB/s)
R1 active weights, FP8 (37 GB)7.7 ms11.0 ms
Dense 70B weights, FP8 (70.6 GB)14.7 ms21.1 ms
R1 active weights, INT4 (18.5 GB)3.9 ms5.5 ms

Source: arithmetic โ€” registry active-parameter count ร— bytes/param รท manufacturer bandwidth (H200 4.8 TB/s, H100 3.35 TB/s per GPU specs); run 2026-10-09. CALCULATED_ESTIMATE: weights-streaming floor only, batch 1.

Two readings. First, on the same GPU at the same precision, R1 pays ~1.9ร— less weight-read time per token than a dense 70B (7.7 vs 14.7 ms on H200) โ€” that gap is MoE's per-token bandwidth advantage in one number. Second, halving weights with precision halves the read: INT4 active reads (18.5 GB) run 3.9 ms on H200, so quantization does double duty for MoE โ€” it shrinks the resident-capacity bill and the per-token stream, at the INT4 quality trade-off the quantization explainer documents.

No throughput claims follow from this floor, deliberately: the cost article's stance is that multi-GPU R1 throughput is PROJECTED territory unless it comes from a real cluster record, and nothing here changes that.

What changes at batch > 1

The two-budget model is batch-dependent, and it's worth being explicit:

  • Dense: each forward step reads the full weight set once regardless of batch size โ€” more tokens per step amortize the same read.
  • MoE at batch 1: reads โ‰ˆ active weights per token (the floor above).
  • MoE at large batch: with enough concurrent tokens, nearly every expert is hit in each step โ€” per-step reads approach the total weight set, and MoE's per-token bandwidth advantage shrinks toward the dense amortization curve. The resident-capacity bill never changes: all experts were resident at step one.

This is why the site's DeepSeek pages treat MoE decode as memory-bandwidth bound (cost article) and why capacity โ€” not activity โ€” drives cluster sizing (minimum viable cluster). The one throughput figure the site publishes (100 tok/s, from its own cluster record) belongs to that page, not to this arithmetic.

FAQ

If only 37B parameters are active, why can't I run DeepSeek R1 on a 2ร—H200 card's worth of memory? Because routing is per token, not per lifetime: over a conversation, every expert fires, and an expert that isn't resident can't fire at all. INT4 gets the resident bill to 372.3 GB at 4K โ€” still multi-GPU โ€” and to 417 GB at 128K.

Is a bigger active ratio worse? It moves the model toward dense behavior: more weight-read per token (the bandwidth table) but fewer total weights for the same active compute. The ratio is a design trade, not a quality dial โ€” nothing in this article's arithmetic says which routing choice is better.

Does MoE change anything about quantization? No โ€” bytes/param is bytes/param. FP8 halves, INT4 quarters both the resident capacity (capacity table) and the per-token read (bandwidth table). The quality trade-offs are the quantization guide's to make: FP16 vs FP8 vs INT4.

Which of these can I actually run locally? Per the capacity table: GLM-4 Flash INT4 fits an 8 GB card; Mistral Large and Ling 3.0 Flash INT4 fit 80 GB; Mixtral 8x22B INT4 is a marginal 80 GB fit (79.5 GB). Full sizing across context lengths: VRAM tiers and the calculator.

Limitations and assumptions

  • Registry figures, arithmetic propagation. Total/active parameters are the OpenGPU Radar model registry's values as published on each model page; weights = parameters ร— bytes/param; percentages = active รท total. Where a registry active figure looks surprising (GLM-4 Flash lists active = total), we publish the registry's numbers rather than our own guesses.
  • Capacity table conditions: batch 1, single GPU, 4,096-token context, run 2026-10-09; totals = weights + KV + 1.2 GB runtime + 0.4 GB activation + 10% margin โ€” identical methodology to the VRAM tiers article. KV confidence: exact GQA (high) for the dense anchor; MLA modeled estimate (medium) for DeepSeek; estimated models (medium) for Mixtral, Nemotron, Mistral Large, Ling and GLM โ€” engine output, only the confidence differs.
  • Bandwidth table: weights-streaming floor only โ€” excludes KV reads, attention, all-reduce, scheduler overhead, and speculative decoding; manufacturer bandwidth specs; batch 1. No tok/s, no benchmarks, no prices.
  • Not covered: training MoE models, expert parallelism, router loss/load-balancing mechanics, and provider-level pricing (rates live on the GPU pages).

Related resources