Dense and Mixture-of-Experts Architectures
What separates dense from mixture-of-experts, what follows from it, the nondeterminism consequence, and when it should change a decision.
Explain the architectural split that now determines most of what you can predict about a model’s cost, latency, and hardware footprint from its name alone.
30.1 The Distinction #
In a dense transformer, every parameter participates in the processing of every token. A 70B dense model performs roughly 70 billion parameters’ worth of arithmetic per generated token.
In a mixture-of-experts model, the feed-forward block at each layer is replaced by a set of expert networks together with a small router network. The router selects a few experts per token, commonly two of eight or eight of many more, and only those run. Total parameters stay large; active parameters per token are a fraction.
The naming convention has become standard, and it encodes the split directly. A model labeled 35B-A3B holds 35 billion total parameters and activates 3 billion per token.66
That single label carries two first-order estimates: total parameters bound what must fit in memory, and active parameters bound the compute per token.
30.2 What Follows From It #
| Property | Governed by | Why |
|---|---|---|
| Memory to hold the model | Total parameters | Every expert must be resident to be routable |
| Arithmetic per token | Active parameters | Only selected experts run |
| Generation speed | Active parameters, first-order only | Decoding is bandwidth bound, and attention, dispatch and offload also bill (§30.6) |
| Quality ceiling | Closer to total | Capacity lives in the full parameter count |
| Batching efficiency | Worse than dense | Different tokens route to different experts |
The consequence that matters is this: MoE decouples what a model knows from what it costs to run. A 35B-A3B model needs the memory of a 35B model and does roughly a 3B model’s arithmetic per token. Active parameters are a first-order proxy for compute, not a speed guarantee: attention, expert dispatch, memory bandwidth, batching and any offload all sit between that count and the tokens per second you observe, so expect a 35B-A3B model to land somewhere between a 3B and a 35B dense model rather than at the 3B figure (§30.6). For anyone running models on their own hardware, that trade is still the reason local inference became practical (§37).
The batching caveat is the underdiscussed case. Dense models batch cleanly because every request exercises the same weights. In an MoE, tokens in the same batch route to different experts, so the serving stack must either gather across experts or accept idle capacity. This is part of why MoE models can show excellent single-stream latency and less impressive aggregate throughput, and why hosted MoE pricing does not always track the active-parameter count as neatly as the arithmetic suggests.
30.3 Why Nearly Everything Is MoE Now #
Dense scaling eventually ran into a wall. Quality tracks parameter count, but serving cost tracks parameter count as well, so every quality gain arrived with a proportional cost increase. MoE breaks the proportionality. Adding experts raises capacity while leaving per-token compute roughly flat.
The practical result is plainly visible in the open-weight landscape. Coding models now ship at sizes like 35B-A3B and 480B-A35B—the latter holding 480 billion parameters and activating 35 billion.66 A 480B dense model would be undeployable outside a large cluster; at 35B active it is merely expensive.
Frontier hosted models are widely believed to be MoE, though the major labs do not publish architecture details. The behavioral evidence is indirect but consistent: pricing that does not scale with apparent capability, and the routing-related nondeterminism described in §5.2.
30.4 The Nondeterminism Consequence #
Section 5.2 listed mixture-of-experts routing as a source of variance without explaining it. The mechanism is now available to us.
The router’s expert selection depends on the token’s hidden state, which is computed in a batch. Batch composition depends on what other traffic the provider is serving at that instant. When two experts have nearly equal router scores, a last-bit floating-point difference flips the selection, a different expert runs, and the output distribution changes.
This is why greedy decoding does not give bit-exact reproducibility against a hosted MoE endpoint under default serving: the batch is not yours to control. It is not an impossibility theorem. Batch-invariant kernels exist—vLLM ships an opt-in mode and lists tested MoE models, DeepSeek and Qwen families among them, while documenting configurations where the guarantee does not yet hold.78 The property is therefore recoverable where you own the serving stack, pin the runtime, and stay inside the tested configurations. What self-hosting buys is the ability to control batching, kernels, sampling, hardware and versions; it does not deliver reproducibility by default. If reproducibility is a hard requirement, that controllability is one of the stronger arguments for local inference, and the configuration work is part of the cost.
30.5 What You Can Infer From a Model Card #
Model cards now carry enough to predict deployment characteristics before you download anything.
Qwen3.6-35B-A3B, Apache 2.0
35B total -> weights must fit in memory
at 4-bit: 35 x 0.5 = ~17.5 GB, plus KV cache
at 16-bit: ~70 GB
3B active -> per-token compute of a ~3B model, an upper bound on
speed rather than a prediction of it (§30.6)
Apache 2.0 -> commercial use permitted, no field-of-use restriction
The arithmetic in §37.3 makes that concrete. The point here is that three fields on a model card—total parameters, active parameters, and license—get you a first pass at deployability without benchmarking anything. They do not settle it: context length and the KV or recurrent state layout decide how much memory the weights leave you, and §30.6 lists what stands between an active-parameter count and observed throughput.
30.6 When the Distinction Should Change a Decision #
Prefer MoE for most local work, and hold the reason at the right altitude: it is the way to get large-model capacity inside memory you own (§37). A 35B-A3B model is not a 3B model wearing a bigger label—attention, expert dispatch, memory bandwidth, batching and any offload all sit between the active-parameter count and the tokens per second you observe—and which architecture wins on quality is workload-dependent rather than settled.
Active parameters predict time-per-token better than total. A large-total, small-active model can feel faster than a smaller dense one.
Size memory on total, size throughput on active, and expect worse batching efficiency than the active count implies.
Hosted MoE adds a variance source that self-hosting lets you control, rather than one it removes for free (§30.4).
A 35B-A3B model scoring near a frontier hosted model on a coding benchmark is a result about capacity, not a claim that three billion parameters are doing frontier work. The capacity is in the 35 billion.
30.7 What This Does Not Change #
Nothing in this section changes how you prompt. Instruction design, context assembly, output contracts, and verification are identical across architectures.
Two things it does change: what you can afford to run yourself, and how much reproducibility you should expect from a hosted endpoint. Both are deployment decisions, and both are the subject of §37.
References cited in this section
2 of 81 · numbering matches the PDF
- 66Qwen team model card for Qwen3.6-35B-A3B together with Unsloth's local-deployment documentation for Qwen3-Coder, https://unsloth.ai/docs/models/tutorials/qwen3-coder-how-to-run-locally, verified September 8, 2026. Vendor documentation and self-reported benchmarks. The parameter counts, Apache 2.0 licensing and reported benchmark figure come from the model card; the Unsloth page is deployment guidance and is not the source of any number here. Cited as product fact for the total-versus-active parameter naming convention, the Apache 2.0 licensing, and the availability of 35B-A3B and 480B-A35B configurations. Note that the card describes a mixed architecture, gated-linear layers interleaved with attention layers, which is why §29.1's KV formula is labeled as full-attention or grouped-query arithmetic rather than as a universal law. The 73.4 percent verified-software-engineering-benchmark figure is the model authors' own reported number, is not independently replicated, and should not be read as a prediction about any particular repository—see reference 55 on why public coding benchmarks travel badly.huggingface.co/Qwen/Qwen3.6-35B-A3B ↗
- 78vLLM documentation: "Batch invariance," "Reproducibility," https://docs.vllm.ai/en/latest/usage/reproducibility/, and "Offline inference," https://docs.vllm.ai/en/latest/serving/offline_inference/, all verified September 25, 2026. Project documentation. Cited for two corrections. Batch-invariant serving is an opt-in mode with an enumerated list of tested models, DeepSeek and Qwen MoE families among them, which is why §30.4 describes hosted nondeterminism as a property of default serving rather than an impossibility in principle; the project also tracks configurations where the guarantee does not yet hold, so the mode is scoped rather than universal, and reproducibility additionally requires pinning kernels, sampling, hardware and versions. And vLLM exposes offline in-process inference through LLM.generate() and LLM.chat(), so it is not API-only (§37.4). Also cited, against the penalty implementation in vllm/model_executor/layers/utils.py, for the §31.2 distinction between the two count-based penalties: the frequency term multiplies an occurrence-count tensor and the presence term a binary occurrence mask, so only the former scales with repetition. An earlier revision of this document described both as scaled by how often the token occurred.docs.vllm.ai/en/latest/features/batch_invariance ↗