Venturing out into the world with AI

Venturing out into the world with AI

🎙 Jon Kleinberg 👥 75K 📅 May 28, 2026 ⏱ 44 min 👁 779 📄 expert opinion 🧭 2026-08-03
Available in: English (current) Français

Keywords

world modelLLMMyhill-Nerodechessnavigation

Summary

Jon Kleinberg, in a talk at the Simons Institute, explores the concept of ‘world models’ in AI systems, particularly large language models (LLMs). He contrasts classical game-playing AI like Deep Blue or AlphaZero, which explicitly represent the game state, with LLMs that generate sequences without an explicit internal state. To probe whether LLMs have world models, he proposes using the Myhill-Nerode theorem from automata theory to induce states from sequences. He illustrates this with examples: a number-choosing game that is isomorphic to tic-tac-toe, and a transformer trained to give driving directions in Manhattan. The latter, when probed, does not reveal a coherent map, suggesting the model lacks a true world model. He discusses the implications for AI agents and the need to understand what LLMs actually learn. The talk is part of a workshop on the role of theoretical computer science in modern machine learning.

145 words

Critical Evaluation

The talk provides a thoughtful and rigorous exploration of a central question in AI: whether LLMs possess world models. Kleinberg’s approach is commendable for grounding the discussion in formal methods, specifically the Myhill-Nerode theorem, which offers a principled way to define and extract states from sequence-generating machines. This is a significant contribution to the interpretability literature, as it moves beyond ad hoc probing to a theoretically motivated framework. The examples are well-chosen and illustrate the concepts clearly. The finding that a transformer trained on Manhattan taxi routes does not yield a coherent map when probed is striking and suggests that LLMs may rely on superficial patterns rather than deep structural understanding. However, the talk is limited by its scope: it focuses on discrete, finite-state domains, and the extension to more complex, continuous real-world scenarios is not addressed. Additionally, the negative result for Manhattan navigation is based on a specific training setup, and it is unclear how generalizable it is. The talk does not engage with alternative approaches to world model extraction, such as causal interventions or representation analysis. The discussion of the Myhill-Nerode theorem is accessible but assumes some familiarity with automata theory. Overall, the talk is intellectually stimulating and offers a fresh perspective, but it leaves many open questions and would benefit from a more comprehensive treatment of the challenges involved. The title is somewhat broad, but the content is well-aligned with the theme of AI venturing into the world. The presence of a short audience question adds value, but the talk would have been stronger with more interactive discussion.

261 words

Title / Content Match

The title is broad but accurately reflects the talk's focus on AI systems interacting with the world and the challenge of understanding their internal models.

Quality & Reliability

8/10

Talk by a leading computer scientist at a prestigious institute, presenting research with clear methodology and references to prior work. However, it is a presentation, not a peer-reviewed publication, and some claims are anecdotal.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk introduces a novel approach to probing world models in LLMs by applying the Myhill-Nerode theorem to induce states from sequences. This provides a formal, theory-grounded method for interpretability, contrasting with ad hoc probing. The empirical finding that a transformer trained on Manhattan navigation does not exhibit a coherent map is a significant contribution to understanding the limitations of current LLMs.

Pour aller plus loin :

97 words

Radar Profile

The radar profile shows high scores in quality and reliability, with moderate scores in quantity and technical level. This indicates a well-founded but specialized talk that may not be accessible to all audiences.

Reliability 8/10