Stanford CS336 Language Modeling from Scratch | Spring 2026 | Lecture 15: Mid/Post-Training

Stanford CS336 Language Modeling from Scratch | Spring 2026 | Lecture 15: Mid/Post-Training

🎙 Percy Liang, Tatsunori Hashimoto 👥 1.2M 📅 May 27, 2026 ⏱ 79 min 👁 19K 📄 lecture 🧭 2026-08-03
Available in: English (current) Français

Keywords

post-trainingSFTRLHFinstruction followingdata quality

Summary

The lecture introduces post-training for language models, focusing on supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF). It contrasts base models like GPT-3 with instruction-following models like ChatGPT, emphasizing the role of post-training in extracting desired behaviors. The talk covers the evolution of SFT data, from FLAN and Self-Instruct to distillation approaches like Alpaca and Vicuna, and more recent agentic data. It highlights the importance of data quality, the challenges of collecting human feedback, and the trade secrets surrounding frontier post-training. The lecture also discusses pitfalls such as data contamination and the need for careful annotation. It concludes with a preview of the next lecture on moving from ChatGPT to reasoning models.

114 words

Critical Evaluation

The lecture provides a comprehensive overview of post-training, specifically SFT and RLHF, from a leading academic perspective. The speakers, Percy Liang and Tatsunori Hashimoto, are well-known researchers in the field, lending credibility to the content. The presentation is well-structured, starting with the motivation for post-training, then detailing the SFT data landscape, and touching on RLHF. The discussion of historical datasets like FLAN and Self-Instruct is valuable for understanding the evolution of techniques. The lecture also candidly addresses the limitations of public knowledge due to trade secrets, which is an important caveat. The technical depth is appropriate for a graduate-level course, assuming familiarity with language models and training. The argumentation is solid, with references to key papers and examples. However, the lecture is not a peer-reviewed source and represents the instructors’ interpretations. The adéquation titre/contenu is excellent, as the title accurately describes the lecture’s focus. Overall, the lecture is a high-quality educational resource, though it may not offer novel research contributions.

160 words

Title / Content Match

Title accurately reflects content: lecture on mid/post-training for language models.

Quality & Reliability

8/10

Lecture by Stanford professors, well-structured, references key papers (RLHF, HH), acknowledges limitations and trade secrets. High credibility but not peer-reviewed.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The lecture provides a structured overview of post-training, synthesizing historical and current approaches. It emphasizes the importance of data quality and the challenges of human annotation. The discussion of trade secrets and the shift towards agentic data offers a contemporary perspective.

Pour aller plus loin :

94 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-balanced lecture with substantial information, high technical depth, and strong reliability. The lecture excels in providing a comprehensive overview of post-training, though it may not offer novel research contributions.

Reliability 8/10