Sherry Yang: Learning World Models and Agents for High-Cost Environments

Sherry Yang: Learning World Models and Agents for High-Cost Environments

🎙 Sherry Yang 👥 845 📅 August 26, 2026 ⏱ 58 min 👁 2 📄 expert opinion 🧭 2026-08-26
Available in: English (current) Français

Keywords

world modelshigh-cost environmentspolicy evaluationreinforcement learninggenerative models

Summary

Sherry Yang’s talk addresses the challenge of applying AI agents to environments where interactions are expensive, such as robotics, ML engineering, and scientific experimentation. She argues that while superhuman performance is achievable in low-cost simulated environments (e.g., AlphaGo, coding competitions), real-world applications are bottlenecked by high interaction costs. The talk outlines three strategies: (1) using learned world models as high-fidelity simulators for robotics, enabling policy evaluation and training without physical hardware; (2) adapting reinforcement learning to handle long action delays in ML engineering; and (3) using compositional generative models to navigate hypothesis spaces in science. She presents concrete examples from her research, including the UniSim model trained on internet-scale video data, and the WorldGym framework for policy evaluation. The talk highlights the potential of world models to reduce costs and enable out-of-distribution testing, but also notes limitations such as hallucination and the loss of internet knowledge in fine-tuned policies.

149 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into a cutting-edge research area, presenting concrete methods and results. The argumentation is solid, building from the problem of high-cost interactions to specific solutions, each supported by examples and quantitative comparisons (e.g., success rates in policy evaluation). The speaker effectively demonstrates the utility of world models for policy evaluation, including the ability to test out-of-distribution scenarios via image editing. The discussion of limitations, such as hallucination and policy overfitting, adds nuance. However, the talk is more of a research overview than a deep dive into any single method, and some claims would benefit from more detailed experimental evidence.

Scientific Rigor, Source Quality, Title Accuracy

The speaker cites her own and collaborators’ papers, including those published at ICLR and on arXiv. The methods are presented with sufficient clarity for an expert audience, but the talk does not provide a systematic comparison with alternative approaches. The title accurately reflects the content. No external sources are cited beyond the speaker’s own work, which is appropriate for a research talk. The talk does not include a public Q&A session, so no audience feedback is available.

195 words

Title / Content Match

The title accurately reflects the content: the talk focuses on learning world models and agents specifically for high-cost environments, as presented.

Quality & Reliability

8/10

Talk by a leading researcher (NYU/DeepMind) presenting peer-reviewed and recent work (ICLR best paper, arXiv preprints). Methods are clearly explained, but the talk is a research overview rather than a systematic review, and some claims (e.g., success rates) are based on specific experiments without full methodological detail.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk presents a coherent research agenda for addressing high-cost environments in AI. The key novelty is the use of large-scale world models trained on internet video data as simulators for robotics, enabling low-cost policy evaluation and training. The WorldGym framework is a concrete contribution, demonstrating that world models can preserve relative policy performance and enable out-of-distribution testing. The talk also discusses adaptations of RL for long action delays and the use of compositional generative models for scientific discovery, though these are less detailed.

Pour aller plus loin :

  • World Models — Background on the concept of world models in AI.
  • Diffusion Transformers — The architecture underlying modern video generation models.
  • Reinforcement Learning — Core algorithm family for training agents.
  • AI for Science — Overview of AI applications in scientific discovery.

131 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and technical level, with a slightly lower but still strong reliability score. This indicates a technically dense and informative talk, though the reliability is slightly tempered by the lack of external validation and the reliance on the speaker's own research.

Reliability 8/10