Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL

Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL

🎙 Aakanksha Chowdhery 👥 1.2M 📅 August 3, 2026 ⏱ 72 min 👁 934 📄 lecture 🧭 2026-08-04
Available in: English (current) Français

Keywords

STaRDeepSeekMathDAPOAIMEreinforcement learning

Summary

This lecture from Stanford’s CS329A course, taught by Aakanksha Chowdhery, focuses on train-time scaling and scaling reinforcement learning for AI agents. The instructor introduces the concept of train-time scaling as a complement to test-time scaling, where models are improved by training on their own filtered outputs. The lecture covers three key papers: STaR (Self-Taught Reasoner), which bootstraps reasoning chains through rationalization and filtering by correctness; DeepSeekMath, which introduces Group Relative Policy Optimization (GRPO) and shows gains from curated math data; and DAPO, which addresses entropy collapse and training instability in long chain-of-thought reasoning. Using the AIME benchmark, the lecture demonstrates how these methods enable smaller models to match or exceed the performance of much larger models. The instructor emphasizes the importance of verifiability in domains like math and code for effective train-time scaling, and discusses the trade-offs between train-time and test-time compute. The lecture concludes with open questions about why majority-at-K accuracy improves while pass-at-K does not, and highlights the practical challenges of implementing reinforcement learning at scale.

168 words

Critical Evaluation

The lecture provides a comprehensive and rigorous overview of train-time scaling techniques, grounded in three influential papers. The instructor, Aakanksha Chowdhery, is a recognized expert with hands-on experience in training large models, lending credibility to the presentation. The content is well-structured, starting with motivation via the AIME benchmark, then detailing each paper’s methodology and contributions, and finally discussing broader implications and open questions. The technical depth is high, with clear explanations of concepts like GRPO, entropy collapse, and dynamic sampling, making it suitable for an audience with some background in machine learning. The lecture also includes a valuable discussion on the trade-offs between train-time and test-time compute, and the importance of verifiability in domains like math and code. However, the lecture is a single perspective and does not critically evaluate the limitations of the presented methods beyond what is mentioned in the papers. The instructor also acknowledges that some plots may be misleading, which adds a layer of nuance. Overall, the lecture is a valuable resource for understanding state-of-the-art techniques in self-improving AI agents, with a strong emphasis on practical implementation details. The adéquation between title and content is excellent, as the lecture directly addresses train-time scaling and scaling RL. The sources cited are the papers themselves, which are peer-reviewed and well-regarded in the field. The lecture does not include any advertising or sponsored content. The audience appears to be graduate students or researchers, but the content is presented clearly enough for a broader technical audience. The discussion of open questions encourages further exploration and critical thinking. Overall, this is a high-quality lecture that provides both theoretical insights and practical guidance.

271 words

Title / Content Match

The title accurately reflects the content, which focuses on train-time scaling and scaling reinforcement learning.

Quality & Reliability

8/10

Lecture from Stanford University by an expert in the field, covering peer-reviewed papers (STaR, DeepSeekMath, DAPO) with technical depth. The content is well-structured and based on established research, though it is a single perspective and not peer-reviewed itself.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

This lecture provides a clear synthesis of three key papers on train-time scaling, highlighting how smaller models can achieve competitive performance through clever training techniques. It emphasizes the importance of verifiability in domains like math and code, and discusses practical implementation details for reinforcement learning. The lecture also raises open questions about the behavior of scaling laws, such as the discrepancy between majority-at-K and pass-at-K accuracy.

Pour aller plus loin :

127 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded lecture with substantial information, high technical depth, and strong reliability. The balance between quantity and quality of information is particularly notable, making it a valuable resource for understanding train-time scaling.

Reliability 8/10