MoE-Study β€” Dense vs. Mixture-of-Experts, matched active parameters

Two decoder-only language models trained from scratch under identical conditions, differing in exactly one thing: whether the feed-forward block is a dense MLP or a sparse top-2-of-4 MoE.

Both checkpoints live in this one repo:

Subfolder Model Total params Active params/token
dense/ Dense FFN 150.1M 150.1M
moe/ Top-2-of-4 MoE 206.8M ~150.1M

The MoE's active-parameter count matches Dense by construction β€” 2 of 4 experts at half the hidden size means identical compute per token. The MoE only spends more memory for extra capacity.

Full write-up, training code, and evaluation harness: github.com/OliverSundaram/MoE-Study


⚠️ These are research artifacts, not usable models

Read this before downloading.

  • Trained for one epoch on ~40.7M tokens β€” neither model is close to converged.
  • WikiText word perplexity is 551 (Dense) and 1,378 (MoE). Generations are largely incoherent.
  • 0.0% on LAMBADA for both β€” at the task floor.
  • No instruction tuning, no RLHF, no safety filtering of any kind.

They exist to answer one narrow question: at matched active compute and matched budget, does sparsity help? They are not fit for any downstream use.


Getting the weights

These are a custom architecture, not a variant of an existing one. The modeling code is not included here, so from_pretrained on this repo alone will not build the model.

Clone the GitHub repo β€” it carries the model definition and loading instructions, and points back at these subfolders for the weights.


Model details

Shared architecture

Both models are the same custom decoder-only transformer:

Layers 12
Attention heads 12
Embedding dim 768
Context length 1024
Vocabulary 50,257 (GPT-2 tokenizer)
Attention Multi-Query β€” one shared K/V projection across all query heads
Normalization Custom pre-norm (learned scale + shift)
Position embeddings Learned absolute
Weight tying None β€” separate input embedding and output head

The one difference

dense/ moe/
FFN block 2-layer GELU MLP 4 experts, top-2 routed
hidden_dim 3072 1536 (per expert)
Router β€” linear β†’ softmax β†’ top-2, renormalized
Aux loss β€” load-balancing term, summed over all 12 layers

Both models share the same unmodified GPT-2 tokenizer, stored once at the repo root.


Training

Identical for both models. Single consumer GPU, no cloud.

Setting Value
Data nampdn-ai/tiny-textbooks
Tokens 39,717 chunks Γ— 1024 = ~40.67M
Epochs 1 (19,858 steps)
Batch size 2 Γ— grad accum 4 = effective 8
Optimizer AdamW, lr 3e-4, weight decay 0.1 (no decay on 1-D params)
Schedule OneCycleLR, cosine, 3% warmup
Grad clipping max-norm 1.0
Precision AMP autocast + GradScaler
Seed 42
Hardware 1Γ— NVIDIA RTX 4060, 8 GB VRAM
Wall-clock ~44.6 min (Dense) Β· ~59.8 min (MoE)

Final losses

Dense MoE
Train loss (final step) 5.166 5.936
Test loss (pure LM) 5.063 5.911
Test loss (+ unscaled aux) n/a 17.91

Dense has the lower loss at every checkpoint.


Evaluation

All benchmarks via lm-evaluation-harness on the final checkpoints.

Benchmark Shots Metric Dense MoE abs(Ξ”) Winner
ARC-Easy 0 acc 29.2% 27.4% 1.8 πŸ”΅ Dense
PIQA 0 acc 55.0% 54.1% 0.9 πŸ”΅ Dense
WikiText 0 word_perplexity 551.0 1,377.8 826.8 πŸ”΅ Dense
LAMBADA (OpenAI) 0 acc 0.0% 0.0% 0.0 βšͺ Tie
WinoGrande 5 acc 50.2% 50.7% 0.5 🟠 MoE
HellaSwag 10 acc_norm 24.9% 25.1% 0.2 🟠 MoE
ARC-Challenge 25 acc_norm 22.9% 23.0% 0.1 🟠 MoE

How to read this:

  • Dense wins on everything sensitive to raw LLM quality β€” perplexity, ARC-Easy, PIQA.
  • WinoGrande, HellaSwag, and ARC-Challenge are won by MoE, but with such a negligible difference, that they could be considered to have an equal accuracy

Inference speed

Greedy decoding, 32-token prompt β†’ 64 new tokens, 5 trials, 2 warmup, no KV cache.

Model Tokens/sec Total params Active params/token
Dense 106.49 Β± 0.30 150.1M 150.1M
MoE 34.40 Β± 0.08 206.8M ~150.1M

MoE is ~3.1Γ— slower despite matched active compute β€” an artifact of unoptimized expert dispatch, not a property of the architecture.

Benchmark charts

ARC-Easy PIQA WikiText LAMBADA WinoGrande HellaSwag ARC-Challenge Speed


Findings

1. Dense won every metric that wasn't already at chance. Most clearly on WikiText perplexity β€” 551 vs 1,378, a 2.5Γ— gap.

2. The routing math is correct. Active parameters match Dense almost exactly. Matched active compute simply didn't buy matched quality at this budget.

3. Routing stayed balanced. The load-balancing term sat on its theoretical floor, so the MoE's gap is not explained by experts collapsing onto each other.

4. Extra capacity needs extra tokens. The MoE has 38% more parameters but saw the same ~40.7M tokens β€” likely far too few to train 4 experts per layer, each seeing only a routed fraction of the stream.

Citation

@misc{sundaram2026moestudy,
  author = {Sundaram, Oliver},
  title  = {MoE-Study: Dense vs. Mixture-of-Experts at Matched Active Parameters},
  year   = {2026},
  url    = {https://github.com/OliverSundaram/MoE-Study}
}

Acknowledgments

License

MIT

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Dataset used to train OliverSundaram/MoE-Study