Data/ML/AI

Mixture of Experts

Also written as MoE, Sparse Model

A model design where the network is split into many specialist sub-networks ('experts') and only a few of them are switched on for each word it processes. This lets a model hold a very large amount of knowledge while costing much less to run than a model of the same size that uses everything every time.

Think of it like

A hospital full of specialists where each patient only sees the two or three doctors they need, instead of every doctor in the building.

Junior or senior?

Most engineers only know MoE as a fact about the models they call through an API.

Senior sounds like

Has trained or served an MoE model and can talk about keeping the load balanced across experts or fitting them across GPUs.

Ask them

“Did you work on the MoE model itself, or use one? If you served it, what made it harder than a normal model?”