Unlocking Adaptive Generative Decision-Making with Diffusion Models

Unlocking Adaptive Generative Decision-Making with Diffusion Models

🎙 Minshuo Chen 👥 2K 📅 July 4, 2026 ⏱ 56 min 👁 50 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

diffusion modelsreinforcement learningreward steeringinference latencysemi-Markov decision process

Summary

The talk, presented by Minshuo Chen at the One World Theoretical Machine Learning seminar, addresses two fundamental challenges in using diffusion models for real-time decision-making. First, it introduces a training-free reward steering framework based on the Doob h-transform, which adapts a pre-trained diffusion policy at inference time to generate high-reward actions without requiring reward differentiability. The method provides convergence guarantees and achieves state-of-the-art empirical performance with modest additional sampling cost. Second, it tackles the issue of inference latency inherent in iterative denoising by formulating latency-aware inference scheduling as a semi-Markov decision process. A meta-controller decides when to invoke the diffusion policy, accounting for delayed execution and action buffering, with a learning algorithm that has theoretical guarantees. The talk emphasizes that diffusion models are not just expressive policy parameterizations but controllable components within sequential decision-making systems.

135 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into two novel approaches for integrating diffusion models into decision-making. The reward steering method is innovative, leveraging the Doob h-transform to avoid training and reward differentiability, which is a practical advantage. The theoretical convergence guarantees add rigor. The latency-aware meta-control addresses a critical practical issue often overlooked, framing it as a semi-Markov decision process and providing a learning algorithm with guarantees. The argumentation is clear and well-structured, with intuitive explanations and examples. However, the talk is a seminar, so some details are omitted, and the empirical results are mentioned but not deeply analyzed.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, with clear mathematical formulations and references to classical concepts like the Doob h-transform and score-based diffusion models. The sources are not explicitly cited in the talk, but the methods are based on established literature. The title accurately reflects the content, covering both reward steering and latency-aware meta-control. The talk is well-organized and the technical level is high, suitable for a specialized audience.

180 words

Title / Content Match

The title accurately reflects the content, focusing on adaptive generative decision-making with diffusion models, covering reward steering and latency-aware meta-control.

Quality & Reliability

8/10

The talk is a research seminar by an academic expert, presenting two principled methods with theoretical guarantees and empirical results. The content is rigorous, but as a seminar, it may not include full peer-reviewed details.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk presents two novel contributions: a training-free reward steering method using the Doob h-transform, and a latency-aware meta-control framework for diffusion policies. The reward steering method is particularly innovative as it avoids the need for reward differentiability and additional training, making it practical for real-world applications. The latency-aware scheduling addresses a critical issue in deploying diffusion policies in real-time systems, providing a principled formulation and algorithm.

Pour aller plus loin :

  • Doob h-transform — The h-transform is a classical tool for conditioning stochastic processes, foundational to the reward steering method.
  • Score-based generative models — The talk builds on score-based diffusion models, and this paper provides the theoretical foundation.
  • Semi-Markov decision process — The latency-aware scheduling is formulated as a semi-Markov decision process, a generalization of MDPs with sojourn times.

130 words

Radar Profile

The radar profile shows high scores in technical level and information quality, with slightly lower scores in quantity and reliability, reflecting the seminar format's depth but limited breadth and lack of peer-reviewed sources.

Reliability 8/10