← Back

MoM: Mixture of MoEs routing with receipts

Mom (Carol Miller) from Futurama, owner of MomCorp.

A dense LLM is a neural net that runs every weight (uses every "neuron" in the neural network) for every token. A MoE (Mixture of Experts) model skips most of them: a router picks a small subset of expert FFNs (feed-forward networks) per token, usually 5 to 30% of total parameters, with newer models trending sparser (DeepSeek-V3 activates ~5%, Mixtral 8x7B activates ~28%: 13B of 47B per token).

Tokens A and B both route to paired experts while the other four experts stay inactive.

We propose MoM (Mixture of MoEs) as a CFR (Cross Family Router) with DaD (Debate-adjudicated Distillation).

The CFR sends the same request to several model families, records their proposals and traces, runs challenge and rebuttal rounds, then adjudicates the debate into weights. That export step is DaD: the accepted debate record can become distillation signal instead of treating the call as a single vote.

CFR sits in the model-router lineage. DaD combines multi-agent debate with multi-teacher distillation.

Single-model MoE still stays inside one checkpoint. Mamba (state-space model) and Jamba (attention plus SSM hybrid) change how one model processes sequences; they do not choose between model families for a request.

Three families argue the next span. The adjudicator writes debate weights into the distillation queue.

The open design question is how large each debate unit should be. The system can compare one token at a time, a short span, or a whole transcript. Those choices have different cost and latency profiles, so they should be explicit modes rather than hidden behavior. Tree of Thoughts frames thought-unit size as the same kind of design choice.

MoM is recursive. A MoM is a Mixture of MoEs and a Mixture of MoMs. Inner MoMs sit underneath outer MoMs. The recursion stops when the receipt budget runs out, regardless of how deep the tree goes.

Root call into nested MoMs. Budget bounds the depth.

MoM needs to know what is already loaded: model hash, adapter hash, KV (key-value) cache digest, allowed tools, and replay key. A loaded model can answer immediately; an unloaded one has to be prepared before it can join the route. If the manifest is wrong, MoM is confidently wrong faster.

Hot mom picks resident FFN experts across multiple MoEs. The cold ones stay off the route.

Geoffrey Hinton put the other mom model on stage at Ai4: superintelligent AI as mother, humans as baby. He told CNN, "If it's not going to parent me, it's going to replace me."

MoA (Mixture of Agents) layers model outputs and aggregates them. MoM is one routed transcript with receipts at every step. Without receipts MoM is MoA with extra hops. A good MoM records the PTA (Proposal, Trace, Adjudication). Then the route can be replayed, scored, and distilled. A bad one leaves you with confidence and no receipt.

Our browser native WebGPU inference runtime, Doppler, can already carry lower-level receipts through its transcript-root contract. The full MoM layer is still a draft surface on the roadmap. Until then: Happy Mother's Day!

Sources

  1. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
  2. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
  3. Mixtral of Experts
  4. Mixture-of-Agents Enhances Large Language Model Capabilities
  5. Mamba: Linear-Time Sequence Modeling with Selective State Spaces
  6. Jamba: A Hybrid Transformer-Mamba Language Model
  7. Tree of Thoughts: Deliberate Problem Solving with Large Language Models
  8. RouteLLM: Learning to Route LLMs with Preference Data
  9. Improving Factuality and Reasoning in Language Models through Multiagent Debate
  10. A Survey on Knowledge Distillation of Large Language Models (multi-teacher)
  11. CNN on Geoffrey Hinton's AI mothers proposal