Mixture of Experts
MoE
- Pronunciation
- MIKS-cher uv EK-spurts
In short
A mixture of experts (MoE) is a neural network design that sends each input to only a few of many small experts, so a huge model costs far less to run.
What is a mixture of experts?
In an ordinary dense model, every parameter takes part in processing every token. In a mixture-of-experts model, some layers contain many parallel expert networks instead of one. A small router network looks at each token and picks the few experts best suited to it, and only those run; the rest of the experts stay idle for that token.
This separates the model's total size from the work done per token. A model can hold hundreds of billions of parameters of knowledge while each token only passes through a fraction of them. Mistral's Mixtral 8x7B, released in 2023, has eight experts in each such layer and uses two for every token, and several of today's largest language models use the same idea.
The idea goes back to a 1991 paper on adaptive mixtures of local experts, and sparse versions for large networks were shown in 2017. MoE models train and answer faster than dense models of the same total size, but they need memory for all the experts, and training has to keep the router from sending everything to a few favorites.
A common misconception is that each expert specializes in a human topic, such as one expert for law and one for code. In practice the router learns its own patterns, often based on token types or syntax, and the specialization is rarely that tidy or easy to name.
Key takeaways
- MoE layers hold many expert networks and a router that picks a few per token.
- Only the chosen experts run, so compute per token stays small.
- Total parameters can be huge while active parameters stay modest.
- All experts must still fit in memory, and routing must stay balanced.
- Experts learn their own patterns, not neat human topics.
Readers ask
What are active parameters?
The parameters actually used for one token. In an MoE model they are much fewer than the total, because only the chosen experts run. Speed and cost follow the active count, while memory follows the total.
Is a mixture of experts several separate models?
No. It is one model in which some layers contain several expert sub-networks. The experts are trained together with the router and the shared layers.
Why use mixture of experts?
To get the knowledge of a very large model at the running cost of a much smaller one. It lets labs grow model capacity without growing the compute needed for every token at the same rate.
See also
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- Neural NetworkAI & Machine Learning, p. 33A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
- Model ParametersAI & Machine Learning, p. 30Model parameters are the internal numbers, such as weights and biases, that a machine learning model learns in training and uses to turn inputs into outputs.
- TransformerAI & Machine Learning, p. 49A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.
- InferenceAI & Machine Learning, p. 24Inference is the stage where a trained machine learning model is used to make predictions or generate output from new data, without changing what it learned.
- Deep LearningAI & Machine Learning, p. 14Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
Spotted a mistake or something missing on this page?Suggest an edit