Systematicity in language models' knowledge and self-knowledge

Systematicity in language models' knowledge and self-knowledge

🎙 Jacob Andreas 👥 305 📅 January 22, 2026 ⏱ 87 min 👁 80 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

language modelsknowledgebeliefssystematicityself-knowledge

Summary

Jacob Andreas presents a seminar on systematicity in language models’ knowledge and self-knowledge. He begins by illustrating failures of current LMs, such as factual errors and inconsistent confidence estimates, and argues these stem from poor internal world models and self-models. He contrasts two philosophical views: interpretationism (behavioral) and representationalism (internal states). He reviews evidence for internal world models, including behavioral benchmarks and probing studies that show linear decodability of world state and factual knowledge from LM representations. He then discusses training objectives that optimize for internal systematicity rather than predictive accuracy, such as deductive closure training, which can improve generalization and self-explanation reliability. The talk concludes with open questions about the nature of LM knowledge and the implications for training trustworthy AI.

122 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the debate on whether language models possess knowledge and beliefs, and proposes a concrete research direction. The argumentation is solid, building from observed failures to a philosophical framework and then to empirical evidence and proposed solutions. The speaker acknowledges limitations and alternative views, making the argument nuanced.

Scientific Rigor, Source Quality, Title Accuracy

The talk references several papers, including ‘Deductive closure training of language models’ (Akyürek et al., 2024) and ‘Training Language Models to Explain Their Own Computations’ (Li et al., 2025), which are relevant and credible. The title accurately reflects the content, focusing on systematicity in knowledge and self-knowledge. The presentation is rigorous, with clear methodology and interpretation of results.

126 words

Title / Content Match

The title accurately reflects the content, which explores systematicity in language models' knowledge and self-knowledge.

Quality & Reliability

8/10

The talk is given by a leading researcher (MIT professor) and presents a coherent research agenda with references to published papers. However, it is a seminar presentation, not a peer-reviewed publication, and some claims are based on ongoing work.

Key Moments

Cited Sources

Concurring Sources

Dissenting Sources

  • On the Dangers of Stochastic Parrots — Argues that LMs are just stochastic parrots without true understanding, contrasting with the speaker's view.

Contribution & Novelties

The talk synthesizes recent research on internal world models in LMs and proposes a novel training objective (deductive closure) that optimizes for systematicity. It provides a clear framework for understanding LM knowledge and self-knowledge, and suggests practical interventions.

Pour aller plus loin :

  • Interpretationism — Philosophical view that beliefs are attributed based on behavior.
  • Representationalism — Philosophical view that mental states are representations.
  • Probing classifiers — Technique used to analyze internal representations of neural networks.

75 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical level, indicating a well-balanced and accessible seminar for a technical audience.

Reliability 8/10

💬 No comments were provided for analysis.