Data/ML/AIHigh signalAround since 2015
Model Distillation
Also written as Knowledge Distillation, Teacher-Student
Training a small, cheap model to imitate a large expensive one, so you get most of the quality at a fraction of the cost and latency.
Think of it like
An experienced chef training an apprentice on the specific dishes the restaurant actually serves.
Junior or senior?
Genuinely advanced. Common in teams whose inference costs got serious enough to force the work.
Ask them
“What drove you to distil rather than keep calling the larger model?”