Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 9: Stochastic Dyn. Program

Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 9: Stochastic Dyn. Program

🎙 Prof. Marco Pavone 👥 1.2M 📅 August 12, 2026 ⏱ 77 min 👁 41 📄 lecture 🧭 2026-08-13
Available in: English (current) Français

Keywords

stochastic dynamic programmingMarkov decision processBellman equationvalue iterationpolicy iteration

Summary

This lecture, part of Stanford’s AA203 course, introduces stochastic dynamic programming for optimal control under uncertainty. The instructor, Prof. Marco Pavone, begins by extending the deterministic optimal control problem to include disturbances modeled as random variables, leading to the Markov decision process (MDP) formulation. Key assumptions include the Markov property and risk-neutrality, which enable the derivation of the Bellman equation. The lecture illustrates the theory with two examples: an inventory control problem and the stochastic LQR problem. In the inventory example, the dynamic programming recursion is solved by hand for a small state space, demonstrating the computation of optimal policies. The stochastic LQR example shows that the optimal control law remains linear, but the cost increases by a constant term related to the noise covariance. The lecture then transitions to infinite-horizon MDPs, introducing the discount factor and the stationary Bellman equation. The concept of the Q-function is introduced as a reformulation that facilitates learning-based control when the transition model is unknown. The lecture concludes by previewing value iteration and policy iteration algorithms for solving infinite-horizon MDPs.

176 words

Critical Evaluation

The lecture provides a rigorous and comprehensive introduction to stochastic dynamic programming, a cornerstone of optimal control and reinforcement learning. Prof. Pavone’s expertise is evident in the clear explanations and the careful handling of mathematical details. The content is well-structured, starting with the problem formulation, then deriving the Bellman equation, and illustrating with examples. The inventory control problem is particularly effective in demonstrating the mechanics of dynamic programming, including the handling of constraints and expectations. The stochastic LQR example elegantly shows that the certainty equivalence principle holds, with the optimal policy unchanged and only the cost affected by noise. The discussion of infinite-horizon MDPs and the Q-function is insightful, highlighting the practical importance of the Q-function in model-free settings. The lecture is technically rigorous, with references to Bertsekas’s textbooks for proofs, and the slides are available online. The main limitation is that it is a lecture, not a peer-reviewed source, but the quality is high. The title accurately reflects the content, and the lecture is suitable for advanced students or practitioners with a background in control theory. Overall, this is an excellent educational resource.

184 words

Title / Content Match

The title accurately reflects the content, which focuses on stochastic dynamic programming, including value iteration and policy iteration.

Quality & Reliability

9/10

Lecture by a renowned expert in autonomous systems, with rigorous mathematical derivations and references to established textbooks. The content is well-structured and pedagogically sound, though it is a lecture rather than peer-reviewed research.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

This lecture provides a clear and rigorous introduction to stochastic dynamic programming, bridging the gap between deterministic optimal control and reinforcement learning. The emphasis on the Q-function’s role in model-free settings is particularly valuable for understanding the foundations of RL.

Pour aller plus loin :

66 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-balanced and comprehensive lecture. The technical depth is high, but the clarity of presentation ensures accessibility for an advanced audience.

Reliability 9/10