For most of its history, distillation was an anecdote field. It worked, often spectacularly, and nobody could tell you in advance by how much. Should the teacher be as strong as possible' Folk wisdom said yes; practitioners kept tripping over cases where a stronger teacher produced a worse student. How much data does distillation need' Depends who you asked. Was distilling ever actually cheaper than just training the small model longer' Shrug. The field ran on vibes and ablations, which is a fine way to write papers and a terrifying way to spend ten million dollars on a training run. Meanwhile, right next door, pretraining had undergone exactly the transformation distillation lacked. The Kaplan scaling laws, then Chinchilla, turned 'how big a model should I train, on how much data'' from a matter of taste into a matter of arithmetic. Loss became a predictable function of parameters and tokens. Budgets became optimization problems. The single most consequential number in the industry...
learn more