Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 18: RL Policy Optimization

Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 18: RL Policy Optimization

🎙 Dr. Daniele Gammelli 👥 1.2M 📅 August 13, 2026 ⏱ 80 min 👁 48 📄 lecture 🧭 2026-08-13
Available in: English (current) Français

Keywords

reinforcement learningpolicy gradientREINFORCEactor-criticvariance reduction

Summary

This lecture, part of Stanford’s AA203 course, focuses on policy optimization methods in model-free reinforcement learning. The speaker, Dr. Daniele Gammelli, begins by contrasting policy optimization with value-based methods, highlighting that policies are explicitly parameterized and optimized directly. The lecture derives the policy gradient theorem, starting from the reinforcement learning objective and using the log-derivative trick to obtain an expression that can be estimated via sampling. This leads to the REINFORCE algorithm, which uses Monte Carlo rollouts to estimate the gradient. The speaker emphasizes that REINFORCE is essentially a weighted version of maximum likelihood, where actions leading to higher returns are upweighted. A significant portion of the lecture addresses the high variance of policy gradient estimators, illustrating the issue with a simple example and introducing two variance reduction techniques: subtracting a baseline and using actor-critic methods. The lecture concludes with a brief mention of recent algorithms and applications, and references the companion textbook ‘Principles of Robot Autonomy’.

157 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides a clear and rigorous derivation of policy gradient methods, building from fundamental concepts. The argumentation is solid, with mathematical derivations that are well-explained and intuitive examples that aid understanding. The speaker effectively connects the material to previous lectures and highlights practical considerations such as variance reduction. The content is highly valuable for students and practitioners seeking a deep understanding of policy optimization.

Scientific Rigor, Source Quality, Title Accuracy

The lecture is scientifically rigorous, presenting standard results in reinforcement learning. It references the companion textbook ‘Principles of Robot Autonomy’ and provides links to course materials. The title accurately reflects the content. The speaker’s credentials and affiliation with Stanford lend credibility. No comments were provided for analysis.

127 words

Title / Content Match

The title accurately reflects the content, which focuses on policy optimization methods in reinforcement learning.

Quality & Reliability

9/10

Lecture from a Stanford University course, presented by a researcher with a PhD in machine learning and optimization, covering established reinforcement learning theory. The content is rigorous, well-structured, and aligns with standard academic material.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

This lecture provides a comprehensive and accessible introduction to policy optimization in reinforcement learning, with a strong emphasis on the derivation of the policy gradient and the REINFORCE algorithm. It effectively bridges theory and intuition, making it a valuable resource for learners. The discussion on variance reduction and actor-critic methods is particularly useful for understanding practical implementations.

Pour aller plus loin :

104 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and reliable educational resource. The lecture excels in information quantity and quality, with a strong technical level and high overall reliability.

Reliability 9/10