Interpretability: From Art Towards Science

Interpretability: From Art Towards Science

🎙 Josh Batson 👥 42K 📅 September 1, 2026 ⏱ 43 min 👁 6 📄 expert opinion 🧭 2026-09-01
Available in: English (current) Français

Keywords

interpretabilityneural networksfeaturescircuitsemergence

Summary

Josh Batson presents a talk on interpretability, framing it as a transition from artisanal analysis to a scientific discipline. He introduces four metaphors for understanding models: physics, patterns, programs, and people. He discusses two ‘ghost stories’—the inscrutable kernel and the alien language—as potential failure modes for interpretability. He then illustrates the physics metaphor with a detailed example of a network learning modular arithmetic, showing how Fourier transforms reveal the learned algorithm. The patterns metaphor is explored through examples of feature learning, including a sentiment direction and a ‘Golden Gate Bridge’ feature, demonstrating that models learn meaningful representations. The programs metaphor is illustrated with a case study of addition in a large language model, showing how the model uses a multiscale representation and reuses components. He concludes by highlighting open problems and the need for a theory of emergence, suggesting a path towards a more scientific interpretability.

146 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the current state of interpretability research, with concrete examples from Anthropic’s work. The argumentation is strong, using illustrative examples to support the claim that models exhibit structured, understandable features. However, the talk is more of an expert opinion and a survey of ongoing work rather than a rigorous scientific presentation. The speaker acknowledges the lack of a strong theory and presents open problems, which adds to the credibility but also limits the conclusiveness of the arguments.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous in its presentation, referencing specific research and phenomena, but it does not cite formal sources in the talk itself. The description provides a link to the workshop page, which may contain further references. The title accurately reflects the content, which is a discussion of the scientific status of interpretability. The talk is well-structured and the examples are clearly explained, contributing to its overall rigor.

166 words

Title / Content Match

The title accurately reflects the content, which discusses moving interpretability from ad-hoc analysis towards a more scientific discipline.

Quality & Reliability

8/10

Talk by a leading researcher from Anthropic, presenting concrete examples and open problems, but with limited formal verification and no peer-reviewed sources cited directly.

Key Moments

Cited Sources

Concurring Sources

  • Linear representations of sentiment in large language models — Referenced in the talk as a study showing sentiment is represented linearly.

Contribution & Novelties

The talk contributes to the interpretability discourse by synthesizing current research and framing it within a scientific perspective. It highlights concrete open problems and suggests a path towards a more rigorous understanding of model internals. The examples, such as the addition case study, provide novel insights into how large models implement arithmetic.

Pour aller plus loin :

94 words

Radar Profile

The radar profile shows high scores in information quantity and quality, with a strong technical level, but slightly lower reliability due to the lack of formal citations. This indicates a talk that is rich in content and technically deep, but relies on the speaker's expertise rather than peer-reviewed sources.

Reliability 7/10