Changing Shapes: Interpretability, AI, and the Future of Human Agency

Changing Shapes: Interpretability, AI, and the Future of Human Agency

🎙 Been Kim 👥 42K 📅 August 31, 2026 ⏱ 40 min 👁 20 📄 expert opinion 🧭 2026-08-31
Available in: English (current) Français

Keywords

interpretabilityagentic interpretabilityhuman agencymechanistic interpretabilityAI safety

Summary

In this talk, Been Kim, a researcher at Google DeepMind, reflects on the evolution of interpretability in AI and argues for a new direction to address the challenges posed by increasingly complex AI systems. She begins by expressing concern about the rapid pace of AI development and the potential for human error. She then traces the history of interpretability, from early attempts at inherently interpretable models to feature attribution and concept-based explanations, highlighting the limitations and failures of these approaches, such as the insensitivity of saliency maps. Kim argues that the goal of exhaustive understanding is no longer feasible with large models and agents. She proposes a distinction between ‘micro interpretability’ (focused on safety-critical components) and ‘macro interpretability’ (focused on practical usefulness). She introduces ‘agentic interpretability’ as a method where AI proactively assists human understanding by building mental models, similar to a teacher. She presents a case study where superhuman chess concepts from AlphaZero were taught to grandmasters, demonstrating that AI can help humans learn new concepts. She concludes by discussing ongoing work, including the use of ’neologisms’ (adding new words to the model’s vocabulary) as a potential approach, and emphasizes the importance of maintaining human agency.

197 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the current state and future directions of interpretability. Kim’s argument is well-structured, moving from a historical review to a proposed new paradigm. She effectively uses the chess grandmaster study as evidence that AI can augment human understanding. The discussion of the limitations of feature attribution, including the mathematical impossibility result, is particularly strong. However, the proposal for ‘agentic interpretability’ is somewhat speculative, and the speaker acknowledges that the general implementation is not yet known. The argumentation is persuasive but relies on a personal perspective and a single case study.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, referencing published work and the speaker’s own research. The speaker is a leading expert in the field, which adds to the credibility. The sources are not explicitly cited in the video, but the content is consistent with known literature. The title accurately reflects the content, which discusses the changing nature of interpretability and the need to focus on human agency. The talk is well-structured and the arguments are presented clearly.

185 words

Title / Content Match

The title accurately reflects the content, which discusses the evolution of interpretability and the need for a new 'agentic' approach to maintain human agency.

Quality & Reliability

8/10

High credibility due to the speaker's position at Google DeepMind and the academic context (IPAM workshop). The talk presents a personal perspective but is grounded in published research and includes concrete examples. However, some claims are speculative and lack detailed evidence.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk contributes a novel perspective on interpretability by proposing ‘agentic interpretability’ as a new paradigm focused on proactive AI assistance in human understanding. It also provides a critical review of existing methods and highlights the need to shift from exhaustive understanding to practical usefulness. The chess grandmaster study is a concrete example of AI augmenting human capabilities.

Pour aller plus loin :

93 words

Radar Profile

The radar profile shows high scores in quality of information and reliability, reflecting the speaker's expertise and the academic context. The quantity of information is moderate, as the talk is a perspective piece rather than a comprehensive review. The technical level is high, but accessible to a knowledgeable audience.

Reliability 8/10