Talk by Jacob Andreas (Massachusetts Institute of Technology)

Talk by Jacob Andreas (Massachusetts Institute of Technology)

🎙 Jacob Andreas 👥 75K 📅 June 10, 2026 ⏱ 32 min 👁 1K 📄 expert opinion 🧭 2026-08-06
Available in: English (current) Français

Keywords

language modelsself-knowledgeintrospectioninterpretabilityfine-tuning

Summary

Jacob Andreas, a professor at MIT, presents research on the ability of language models to describe their own internal workings. He begins by demonstrating that models often give incorrect confidence estimates and fail to accurately describe their own policies and algorithms. He contrasts this with human introspection, which is also unreliable, but argues that because models are engineered systems, we might be able to train them to be better at self-reporting. The talk outlines three methods for generating ground-truth explanations: behavioral counterfactuals, feature labeling, and activation patching. These methods are used to create training data for fine-tuning models to answer questions about their own behavior and representations. Results show that fine-tuned models improve at answering these questions, and the approach generalizes to new inputs. The talk concludes by discussing potential applications for interpretability and control, and acknowledges limitations and open questions.

141 words

Critical Evaluation

The talk presents a compelling and well-structured argument for improving language models’ self-knowledge. Andreas demonstrates a clear problem: models are often overconfident and provide inaccurate descriptions of their own decision-making processes. He draws an interesting parallel to human introspection, which is also flawed, but argues that engineered systems offer unique opportunities for intervention. The proposed solutions are methodologically sound, leveraging existing interpretability techniques (counterfactual analysis, feature labeling, activation patching) to generate supervision data. The results, while preliminary, show significant improvements on the targeted tasks. However, the talk is limited in scope: it focuses on simple, toy examples and does not address the scalability of these methods to more complex, real-world scenarios. Additionally, the speaker does not discuss potential risks or ethical considerations of improving models’ self-knowledge, such as the possibility of models providing misleading explanations that are more convincing. The sources cited are relevant and credible, including prior work like the Selfie paper. Overall, the talk is a valuable contribution to the field of interpretability, but it leaves many open questions and does not fully address the broader implications.

179 words

Title / Content Match

The title is generic but accurately reflects the content: a research talk by Jacob Andreas.

Quality & Reliability

8/10

Talk by a leading researcher at MIT, presenting ongoing research with clear methodology and references to prior work. Claims are supported by examples and references to published studies, though some results are not fully detailed.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk presents a novel framework for training language models to better describe their own internal states and behaviors, using interpretability tools as supervision. This is an original contribution to the field of interpretability, as it directly addresses the gap between models’ actual behavior and their self-reports. The approach is practical and could be extended to more complex scenarios.

Pour aller plus loin :

  • Selfie: Self-interpretation of language models — A paper by Carl Vondrick’s group showing zero-shot self-knowledge in models.
  • Activation patching — A technique for identifying causal contributions of internal representations.
  • Interpretability research at MIT — Jacob Andreas’s lab page, with links to related publications.

107 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and informative talk. The strongest aspects are the quantity and quality of information, as well as the technical depth. The reliability is also high, given the speaker's expertise and the use of established methods.

Reliability 8/10