Skip to content

What is…

Distillation

Training a small 'student' model to imitate a big 'teacher' model — most of the smarts at a fraction of the cost.

In distillation, a large, expensive model generates answers (sometimes including its reasoning), and a smaller model is trained to reproduce them. The student learns the teacher's behavior far more efficiently than it could learn from raw data alone.

It's why today's small, cheap, fast models are often better than last year's giant ones — and why labs guard their models' outputs: distilling a competitor's model is a hot-button issue.

💡 Think of it like

A master chef training an apprentice. The apprentice never reads the whole library of cookbooks — they learn by copying the master's dishes.

📌 Example

"Mini", "Flash" and "Haiku"-style models are typically distilled or trained with help from their bigger siblings.

🧠 Test yourself

Which of these describes Distillation?

Related terms

🎮 Learn AI by playing

40 bite-size missions, boss battles and a certificate. Free.

Start the bootcamp →

☀️ One AI term every morning

Plus the day's top stories, in your inbox by 8am.