Diffusion in RL and robotics: how expressive policies changed how we use continuous actions

Diffusion in RL and robotics: how expressive policies changed how we use continuous actions

🎙 Sergey Levine 👥 75K 📅 August 6, 2026 ⏱ 48 min 👁 448 📄 expert opinion 🧭 2026-08-08
Available in: English (current) Français

Keywords

diffusion policyaction chunkrobotic foundation modelsoffline RLflow models

Summary

Sergey Levine presents an overview of how diffusion and flow models have transformed continuous-action policies in reinforcement learning (RL) and robotic control. He begins by illustrating the problem with a video of a robot folding shorts, driven by a 4-billion-parameter model with a DiT-style diffusion output. He explains that traditional Gaussian policies are insufficient when data contains multiple modes, as averaging them leads to poor actions. Diffusion models, by contrast, can represent complex distributions over high-dimensional action spaces. He traces the evolution from the early Diffuser method, which applied unconditional diffusion to trajectories, to the more practical Diffusion Policy, which uses conditional diffusion over action chunks. He highlights the importance of action chunks, noting that they are essential for imitation learning and beneficial for RL. He then discusses robotic foundation models, which combine large-scale pre-trained vision-language models with robot data, and explains how diffusion policies integrate with these models. He concludes by mentioning recent work on large-scale models built on these principles, such as π0, and hints at future directions.

170 words

Critical Evaluation

The talk provides a valuable expert perspective on the application of diffusion models to continuous control in RL and robotics. Levine’s argumentation is clear and grounded in practical experience, supported by concrete examples and references to key papers. He appropriately acknowledges the empirical nature of the benefits, noting that the reasons for the success of action chunks are not fully understood. The presentation is rigorous in its technical explanations, but it is not a systematic review; it is an opinionated overview based on the author’s research and selected works. The sources cited are primarily the author’s own and other prominent papers in the field, which are credible but not exhaustive. The talk does not address potential limitations or failure cases in depth, which could be seen as a weakness. The title accurately reflects the content, and the talk is well-structured, moving from motivation to historical context to current practice. Overall, it is a high-quality, informative talk suitable for an audience with some background in machine learning.

166 words

Title / Content Match

The title accurately reflects the content, which focuses on the role of diffusion models in RL and robotics, particularly for continuous action policies.

Quality & Reliability

8/10

Talk by a leading researcher (Sergey Levine) at a prestigious institute (Simons Institute). Content is based on established research and personal experience, but is not peer-reviewed and represents an expert perspective rather than a systematic review.

Key Moments

Cited Sources

Concurring Sources

  • Diffusion Policy — The paper by Chi et al. that introduced the Diffusion Policy method, which is central to the talk.
  • Diffuser — The paper by Janner et al. that proposed Diffuser, an early diffusion-based planning method.
  • π0: A Vision-Language-Action Flow Model — The paper describing the π0 model, which is an example of a large-scale diffusion-based policy.

Dissenting Sources

  • No discordant sources found — The talk does not present conflicting viewpoints or sources.

Contribution & Novelties

The talk provides a synthesis of recent developments in using diffusion models for continuous control, highlighting the shift from simple Gaussian policies to expressive diffusion-based policies. It emphasizes the practical importance of action chunks and the integration with large-scale pre-trained models. The speaker shares insights from his own research and the broader community, offering a perspective on why these methods work despite theoretical expectations.

Pour aller plus loin :

149 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and informative talk. The high technical level and information quality reflect the speaker's expertise and the depth of content.

Reliability 8/10

💬 No comments were provided for analysis.