[M2L 2025] 3.2 Reinforcement Learning for LLMs - Jessica Hamrick

[M2L 2025] 3.2 Reinforcement Learning for LLMs - Jessica Hamrick

🎙 Jessica Hamrick 👥 3K 📅 November 12, 2025 ⏱ 51 min 👁 234 📄 science communication 🧭 2026-08-15
Available in: English (current) Français

Keywords

RLHFRLVRPPODPOchain-of-thought

Summary

Jessica Hamrick, formerly of DeepMind, presents a comprehensive overview of reinforcement learning (RL) for large language models (LLMs). She begins by contrasting pre-trained LLMs with aligned models, highlighting the alignment problem. She then introduces the RL framework for LLMs, framing prompts as states, tokens as actions, and the model as the policy. She discusses various reward types: RLHF (human preferences), RLVR (verifiable rewards), RLEF (execution feedback), and RLAIF (AI feedback). She explains the RLHF pipeline, including reward model training and PPO, and contrasts it with DPO as a simpler alternative. She then covers RLVR, tracing its origins to chain-of-thought prompting and showing how RL can improve reasoning without extensive SFT data. She also touches on RLEF for agentic tasks like code editing and tool use. Finally, she discusses open research questions, including reward hacking, generalization, and the balance between exploration and exploitation. The talk is well-structured, accessible, and grounded in recent developments like OpenAI’s o1 and DeepSeek-R1.

157 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides a valuable synthesis of RL techniques for LLMs, clearly explaining the conceptual framework and differentiating between reward types. The argumentation is solid, building from foundational concepts to recent advances, and is supported by concrete examples (e.g., o1’s reasoning trace, DeepSeek-R1’s aha moment). The speaker effectively communicates the intuition behind PPO and DPO without oversimplifying, and she highlights practical considerations such as reward hacking and the complexity of online RL. The discussion of open questions adds depth and encourages further exploration.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor by referencing key papers (InstructGPT, DeepSeek-R1, DPO, chain-of-thought) and accurately describing their contributions. The sources are credible and appropriately cited. The title accurately reflects the content, which is a focused overview of RL for LLMs. The presentation is well-organized and technically accurate, though it does not delve into mathematical details, which is appropriate for a summer school lecture.

161 words

Title / Content Match

The title accurately reflects the content, which is a comprehensive overview of reinforcement learning techniques applied to large language models.

Quality & Reliability

8/10

The talk is given by a researcher with extensive experience at DeepMind, providing a clear and accurate overview of RL techniques for LLMs. It covers established methods (RLHF, PPO, DPO) and recent developments (RLVR, RLEF, RLIF) with appropriate references to key papers. The presentation is well-structured and technically sound, though it remains at an introductory level without deep mathematical derivations.

Key Moments

Cited Sources

Concurring Sources

  • OpenAI o1 System Card — Provides details on o1's reasoning and safety evaluations.
  • DeepSeek-R1 Technical Report — Details the RLVR approach and results.

Contribution & Novelties

The talk provides a clear and up-to-date synthesis of RL techniques for LLMs, bridging classical RL concepts with modern LLM training. It effectively categorizes reward types (RLHF, RLVR, RLEF, RLAIF) and explains their applications, making it a valuable resource for researchers and practitioners. The discussion of open questions highlights current challenges in the field.

Pour aller plus loin :

103 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a moderate technical level. This indicates a well-balanced, informative talk that is accessible to a broad audience while maintaining scientific rigor.

Reliability 8/10

💬 Sur les 0 commentaires analysés, aucune tendance n'est disponible.