Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 5 - LLM tuning

Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 5 - LLM tuning

🎙 Afshine Amidi, Shervine Amidi 👥 1.2M 📅 November 14, 2025 ⏱ 107 min 👁 63K 📄 lecture 🧭 2026-08-06
Available in: English (current) Français

Keywords

RLHFPPODPOpreference tuningreward model

Summary

This lecture, part of Stanford’s CME295 course, covers LLM tuning with a focus on aligning models with human preferences. It begins by recapping pre-training and supervised fine-tuning (SFT), then introduces preference tuning as a third step. The lecture explains the need for preference tuning, highlighting the difficulty of creating high-quality SFT datasets and the ability to inject negative signals. It details data collection methods for preference pairs, including pairwise comparisons and LLM-as-a-judge. The core of the lecture is an overview of RLHF (Reinforcement Learning from Human Feedback), starting with RL basics (agent, state, action, policy, reward) and mapping them to LLM concepts. It covers reward modeling, the Bradley-Terry formulation, and then delves into reinforcement learning approaches, particularly PPO (Proximal Policy Optimization) and its variants (clip, KL-penalty). The lecture also discusses challenges, on-policy vs off-policy methods, Best-of-N sampling, and concludes with DPO (Direct Preference Optimization) as an alternative to RLHF. Throughout, the instructors provide clear explanations and examples, making complex topics accessible.

161 words

Critical Evaluation

The lecture provides a solid, structured introduction to LLM tuning, specifically preference tuning and RLHF. The instructors, Afshine and Shervine Amidi, demonstrate deep expertise in the subject, presenting the material in a logical progression from basic concepts to advanced techniques. The value of the information is high for an educational context, as it covers the key methods (RLHF, PPO, DPO) that are foundational to modern LLM alignment. The argumentation is coherent, with clear motivations for each step, such as the limitations of SFT and the benefits of preference tuning. The scientific rigor is evident in the precise definitions and mathematical formulations, such as the Bradley-Terry model and the PPO objective. However, the lecture lacks critical discussion of the limitations and potential pitfalls of these methods, such as reward hacking or the challenges of collecting unbiased preference data. The sources cited are minimal, primarily the course syllabus, which is appropriate for a lecture but limits the ability to verify claims independently. The title accurately reflects the content, and the lecture is well-paced with effective use of examples. Overall, this is a high-quality educational resource, though it could benefit from more critical analysis and references to primary literature.

196 words

Title / Content Match

The title accurately reflects the content, which focuses on LLM tuning techniques, specifically preference tuning and RLHF.

Quality & Reliability

8/10

Lecture from Stanford University, presented by adjunct lecturers with expertise in AI. Content is structured, covers established methods (RLHF, PPO, DPO) with clear explanations. No citations to external sources within the lecture itself, but the syllabus link provides additional resources. The lecture is educational and aligns with current research, though it lacks critical discussion of limitations.

Chapters

Cited Sources

Concurring Sources

  • InstructGPT paper — The lecture's RLHF overview aligns with the methods described in this paper.
  • PPO paper — The lecture's discussion of PPO is consistent with the original PPO algorithm.

Dissenting Sources

Contribution & Novelties

The lecture provides a comprehensive overview of LLM tuning, specifically focusing on preference tuning and RLHF. It offers a clear pedagogical explanation of complex concepts, making them accessible to a broad audience. The inclusion of DPO as an alternative to RLHF is particularly valuable, as it highlights recent advancements in the field.

Pour aller plus loin :

  • RLHF paper (InstructGPT) — Original paper introducing RLHF for LLMs.
  • PPO paper — Proximal Policy Optimization algorithms.
  • DPO paper — Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

88 words

Radar Profile

The radar profile shows high scores in quantity and quality of information, indicating a comprehensive and well-structured lecture. The technical level is moderately high, suitable for an advanced audience, and the reliability is strong due to the academic context.

Reliability 8/10