Keynote- The challenge of building logical, language instructible AI systems

Keynote- The challenge of building logical, language instructible AI systems

🎙 Prof. Csaba Szepesvari 👥 3K 📅 March 3, 2026 ⏱ 25 min 👁 87 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

logical reasoninglanguage instructiblerule followingoff-distribution generalizationverifiers

Summary

In this keynote, Prof. Csaba Szepesvari discusses the challenge of building AI systems that are both logical and can follow natural language instructions. He argues that while large language models have made impressive progress in understanding and generating human language, they still struggle with rule following and logical consistency, especially in novel situations. He introduces the concept of ‘associative retrieval dominating logic’, where models rely on pattern matching from training data rather than true reasoning, leading to failures on simple puzzles when slightly altered. He identifies two main problems: imitation learning (supervised learning) which trains models to replicate data without understanding underlying thoughts, and the difficulty of rule induction, which can be exponentially hard for verifiers. He presents a theoretical result showing that learning to compare numbers can require exponentially many examples for common algorithms. He concludes by suggesting that achieving 100% correctness requires careful problem decomposition or exploiting symmetries, and calls for new approaches to build AI systems that can truly reason and follow rules.

166 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the limitations of current AI systems, particularly the gap between performance on benchmarks and real-world rule following. The argumentation is solid, building from concrete examples (the ‘unpuzzle’ puzzle) to theoretical results (exponential sample complexity for verifiers). The speaker effectively communicates complex ideas in an accessible manner, making a compelling case for the need to address these challenges. The discussion of imitation learning and the lack of ’thinking’ in training data is particularly insightful.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor by referencing specific papers (e.g., ’next token prediction’ by fellow Googlers, and a paper on verifier learning with colleagues). The sources are credible, coming from leading researchers in the field. The title accurately reflects the content, focusing on the challenge of building logical, language-instructible AI systems. The talk does not overstate claims and acknowledges the open nature of the problem. The description provides links to the organization’s website and playlist, which are relevant for further exploration.

175 words

Title / Content Match

The title accurately reflects the content: the talk focuses on the challenge of building AI systems that are both logical and can follow natural language instructions.

Quality & Reliability

8/10

The talk is given by a leading researcher in AI (Google DeepMind, University of Alberta) and presents well-reasoned arguments supported by references to recent research (e.g., 'next token prediction' paper, verifier learning paper). The claims are plausible and align with current debates in the field. However, the talk is a keynote and not a peer-reviewed publication, and some statements are anecdotal (e.g., the 'unpuzzle' example). Overall, high reliability.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk provides a clear articulation of the challenge of building AI systems that are both logical and language-instructible, highlighting the tension between associative retrieval and rule following. It offers a novel perspective on why current models fail on simple puzzles when slightly altered, and presents a theoretical result on the exponential sample complexity of learning verifiers. The talk also emphasizes the importance of off-distribution generalization and the need for new approaches beyond imitation learning.

Pour aller plus loin :

  • Next Token Prediction paper — Discusses the limitations of next token prediction for reasoning.
  • Verifier learning paper — Explores the difficulty of training verifiers for reasoning tasks.
  • Off-distribution generalization — Overview of the challenge of generalizing to new distributions.

119 words

Radar Profile

The radar profile shows high scores in quality of information and reliability, reflecting the speaker's expertise and the soundness of the arguments. The quantity of information is moderate, as the talk is a keynote with a focused scope. The technical level is moderate, accessible to a broad audience. Overall, the profile indicates a well-balanced, credible presentation.

Reliability 8/10

💬 No comments were provided for analysis.