On Policy Distillation - Using LLMs to train better LLMs

On Policy Distillation - Using LLMs to train better LLMs

🎙 Neural Breakdown with AVB 👥 34K 📅 August 23, 2026 ⏱ 30 min 👁 31 📄 science communication 🧭 2026-08-23
Available in: English (current) Français

Keywords

on-policy distillationknowledge distillationLLMtrainingself-distillation

Summary

The video explains on-policy distillation (OPD) and on-policy self-distillation (OPSD), techniques used to train large language models (LLMs) by transferring knowledge from a teacher model to a student model. It contrasts OPD with supervised fine-tuning (SFT) and reinforcement learning (RL), highlighting that OPD provides dense, token-level feedback on trajectories generated by the student itself. The video discusses two modes of teacher supervision: sampled token and full vocabulary distillation, and explains the difference between forward and reverse KL divergence, including their trade-offs. It introduces on-policy self-distillation, where the same model acts as its own teacher using privileged information, and discusses the ‘privileged illusion’ pitfall. Finally, it offers guidance on when to use SFT, DPO, RL, OPD, or OPSD, and mentions practical resources like the Thinking Machines blog and open-source repositories.

129 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides a valuable conceptual framework for understanding on-policy distillation, clearly distinguishing it from other training methods. The argumentation is solid, using intuitive examples and analogies to explain complex concepts like forward vs. reverse KL divergence. The explanation of the ‘privileged illusion’ is particularly insightful, highlighting a subtle but important pitfall. However, the video could benefit from more concrete examples or case studies to strengthen its claims.

Scientific Rigor, Source Quality, Title Accuracy

The video references several key papers and resources, including the Thinking Machines blog and a GitHub repository, which adds credibility. However, it does not provide formal citations or links to the original papers, making it difficult to verify claims. The title accurately reflects the content, and the video stays on topic throughout.

135 words

Title / Content Match

The title accurately reflects the content, which focuses on explaining on-policy distillation and its variants.

Quality & Reliability

7/10

The video provides a clear and structured explanation of on-policy distillation, referencing key papers and concepts. However, it lacks formal citations and relies on anecdotal examples, limiting its scientific rigor.

Chapters

Cited Sources

Concurring Sources

Contribution & Novelties

The video offers a clear and accessible explanation of on-policy distillation, a topic that is gaining prominence in LLM training. It synthesizes recent developments and provides practical guidance on when to use different training methods. The discussion of the ‘privileged illusion’ and the distinction between forward and reverse KL are particularly valuable for practitioners.

Pour aller plus loin :

97 words

Radar Profile

The radar profile shows a balanced performance across all dimensions, with slightly higher scores in information quantity and technical level, indicating a well-rounded educational video.

Reliability 7/10

💬 No comments were provided for analysis.