Lecture 11 (Policy Search) | MIT 6.832 (Underactuated Robotics), Spring 2021

Lecture 11 (Policy Search) | MIT 6.832 (Underactuated Robotics), Spring 2021

🎙 underactuated 👥 17K 📅 March 31, 2021 ⏱ 60 min 👁 1K 📄 lecture 🧭 2026-08-05
Available in: English (current) Français

Keywords

policy searchLQRgradient descentnon-convex optimizationunderactuated robotics

Summary

This lecture from MIT’s Underactuated Robotics course introduces model-based policy search, a method for optimizing control policies directly in the parameter space rather than through dynamic programming or trajectory optimization. The instructor motivates the approach by noting that for many complex systems, simple controllers can be effective even when the dynamics or cost-to-go functions are intractable. The lecture focuses on the linear quadratic regulator (LQR) as a special case, showing that the policy search problem is non-convex but can still be solved via gradient descent without local minima. It also discusses the need to define a distribution over initial conditions to obtain a scalar objective. The instructor then presents a practical algorithm that leverages trajectory optimization techniques to optimize controller parameters. The lecture balances theoretical insights with practical considerations, highlighting recent results and open challenges.

135 words

Critical Evaluation

The lecture provides a solid introduction to policy search, a key topic in reinforcement learning and optimal control. The instructor, Russ Tedrake, is a leading expert in the field, and the content is well-structured and rigorous. The motivation for policy search is clearly articulated: for many real-world systems, the dynamics and cost-to-go functions are too complex to model accurately, yet simple controllers can perform well. The LQR case is analyzed in detail, demonstrating that even though the optimization problem is non-convex, gradient descent can be used effectively because there are no local minima. This is a valuable theoretical result that justifies the use of gradient-based methods. The lecture also highlights the importance of specifying a distribution over initial conditions, which is a crucial practical consideration. The algorithm presented, which adapts trajectory optimization techniques to optimize controller parameters, is a useful tool for practitioners. The lecture is well-paced and includes intuitive explanations, though some parts assume prior knowledge of control theory and linear algebra. The lack of external references is a minor weakness, but the lecture is part of a comprehensive course with extensive notes. Overall, this is a high-quality educational resource that effectively bridges theory and practice.

197 words

Title / Content Match

The title accurately reflects the content: a lecture on policy search methods in the context of underactuated robotics.

Quality & Reliability

8/10

Lecture from MIT OpenCourseWare by a recognized expert in robotics and control. Content is rigorous, with theoretical foundations and practical algorithms. No external sources cited, but the lecture is part of a well-established course.

Key Moments

Contribution & Novelties

The lecture provides a clear and accessible introduction to policy search, a fundamental concept in reinforcement learning and optimal control. It highlights the advantages of searching directly in policy space, especially for systems with complex dynamics. The analysis of LQR policy search, including the non-convexity and the absence of local minima, is a key insight. The lecture also offers a practical algorithm that combines trajectory optimization with policy parameter updates.

Pour aller plus loin :

106 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-balanced lecture with strong information content, technical depth, and reliability. The lecture excels in providing both theoretical foundations and practical algorithms.

Reliability 8/10