TheSequence

TheSequence

The Sequence Knowledge #902: Learning About Distillation: When the Dataset Becomes the Teacher

Once models can generate their own curricula, data stops being a static resource and becomes a transmission medium for intelligence.

Jul 28, 2026
∙ Paid

For most of machine learning history, data was treated as geology. It already existed somewhere in the world—in books, websites, code repositories, conversations, photographs, and databases. The researcher’s job was to excavate it, clean it, tokenize it, and feed it into a model.

Large language models changed this relationship. A capable model is not only a consumer of data. It can produce questions, answers, explanations, critiques, preference labels, tool traces, textbooks, code exercises, and entire miniature curricula.

This creates a new training primitive. Instead of asking an expensive model to answer every production query forever, we ask it to manufacture the experience from which a smaller model learns. The teacher runs offline. Its outputs become a dataset. The dataset trains the student. The teacher disappears at inference time, but some of its behavior remains embedded in the student.

That is synthetic data as distillation.

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Jesus Rodriguez · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture