What is mixture-of-experts?
An architecture that activates only a subset of a model's sub-networks (experts) for each query instead of running the whole model, keeping the model huge while cutting the actual compute per request.
An architecture that activates only a subset of a model's sub-networks (experts) for each query instead of running the whole model, keeping the model huge while cutting the actual compute per request.
Get it in your inbox every day
Daily headlines and summaries, a weekly synthesis every Sunday, and a monthly report at the start of each month. Sent at 8AM KST, which is the evening before in the US.
Free · no ads · one-click unsubscribe