
Stanford CS336 Language Modeling from Scratch | Spring 2026 | Lecture 15: Mid/Post-Training
Keywords
Summary
114 words
Critical Evaluation
The lecture provides a comprehensive overview of post-training, specifically SFT and RLHF, from a leading academic perspective. The speakers, Percy Liang and Tatsunori Hashimoto, are well-known researchers in the field, lending credibility to the content. The presentation is well-structured, starting with the motivation for post-training, then detailing the SFT data landscape, and touching on RLHF. The discussion of historical datasets like FLAN and Self-Instruct is valuable for understanding the evolution of techniques. The lecture also candidly addresses the limitations of public knowledge due to trade secrets, which is an important caveat. The technical depth is appropriate for a graduate-level course, assuming familiarity with language models and training. The argumentation is solid, with references to key papers and examples. However, the lecture is not a peer-reviewed source and represents the instructors’ interpretations. The adéquation titre/contenu is excellent, as the title accurately describes the lecture’s focus. Overall, the lecture is a high-quality educational resource, though it may not offer novel research contributions.
160 words
Title / Content Match
Title accurately reflects content: lecture on mid/post-training for language models.
Quality & Reliability
8/10
Lecture by Stanford professors, well-structured, references key papers (RLHF, HH), acknowledges limitations and trade secrets. High credibility but not peer-reviewed.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
Cited Sources
- CS336 Course Website — Course schedule and syllabus
- CS336 Course Enrollment — Information about enrolling in the course
- Stanford AI Programs — Overview of Stanford's online AI programs
- Course Playlist — Full playlist of lectures
Concurring Sources
- RLHF paper — Detailed recipe for RLHF, referenced in lecture.
- Anthropic's HH paper — Human feedback data collection, referenced in lecture.
Contribution & Novelties
The lecture provides a structured overview of post-training, synthesizing historical and current approaches. It emphasizes the importance of data quality and the challenges of human annotation. The discussion of trade secrets and the shift towards agentic data offers a contemporary perspective.
Pour aller plus loin :
- RLHF paper (Learning to Summarize from Human Feedback) — Foundational paper on RLHF.
- Anthropic’s HH paper — Details on human feedback data collection.
- Self-Instruct paper — Introduces self-instruction for data generation.
- Alpaca paper — Distillation approach for instruction tuning.
- Tulu 3 paper — Recent work on post-training data.
94 words
Radar Profile
The radar profile shows high scores across all dimensions, indicating a well-balanced lecture with substantial information, high technical depth, and strong reliability. The lecture excels in providing a comprehensive overview of post-training, though it may not offer novel research contributions.