What is mixture-of-experts?

An architecture that activates only a subset of a model's sub-networks (experts) for each query instead of running the whole model, keeping the model huge while cutting the actual compute per request.

Briefings mentioning this
2026-09-232026-08-02

Terms seen alongside this one