Epistemological challenges in the study of “Theory of Mind” in LLMs and humans

Epistemological challenges in the study of “Theory of Mind” in LLMs and humans

🎙 Sean Trott 👥 305 📅 October 9, 2025 ⏱ 80 min 👁 62 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

Theory of MindLLMfalse belief taskdistributional statisticsconstruct validity

Summary

Sean Trott presents a seminar on epistemological challenges in studying Theory of Mind (ToM) in large language models (LLMs) and humans. He introduces the concept of using LLMs as ‘model organisms’ to test hypotheses about human cognition, particularly the language exposure hypothesis. He describes a study using the false belief task adapted for LLMs, showing that GPT-3 exhibits sensitivity to implied belief states but performs below human levels. He then discusses a replication with 55 open-weight models, finding that some show sensitivity but none reach human performance. Trott explores three interpretive views: the ‘duck test’, ‘axiomatic rejection’, and ‘differential construct validity’. He argues for the latter, suggesting that tasks may measure different constructs in LLMs and humans. He outlines theoretical and empirical perspectives to evaluate this, including plausibility, face validity, convergent validity, and process mechanisms. The talk emphasizes the need for methodological innovation in both cognitive science and AI.

149 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the use of LLMs as model organisms for studying human cognition, particularly Theory of Mind. Trott presents empirical evidence from his own studies, showing that LLMs exhibit some sensitivity to belief states but fall short of human performance. He carefully discusses limitations, such as data contamination and reproducibility, and proposes a nuanced framework (differential construct validity) for interpreting LLM performance. The argumentation is solid, building from specific experiments to broader epistemological considerations, and he acknowledges the work-in-progress nature of his theoretical contributions.

Scientific Rigor, Source Quality, Title Accuracy

Trott demonstrates scientific rigor by referencing his published papers (Jones et al., 2024; Trott et al., 2023) and discussing methodological details. He addresses potential criticisms of the false belief task and the challenges of studying closed-source models. The title accurately reflects the content, focusing on epistemological challenges. The talk is well-structured and grounded in empirical evidence, though some parts are based on unpublished work, which is clearly indicated.

171 words

Title / Content Match

The title accurately reflects the content, which focuses on epistemological challenges in studying Theory of Mind in LLMs and humans.

Quality & Reliability

8/10

The talk is given by a researcher with relevant expertise, presents empirical studies with published results, and discusses methodological limitations. However, it is a seminar presentation, not a peer-reviewed article, and some claims are based on unpublished work.

Key Moments

Cited Sources

  • Comparing humans and large language models on an Experimental Protocol Inventory for Theory of Mind Evaluation (EPITOME) — Referenced as a publication by Jones, Trott, & Bergen (2024) in the talk.
  • Do large language models know what humans know? — Referenced as a publication by Trott, Jones, Chang, Michaelov, & Bergen (2023) in the talk.

Concurring Sources

Contribution & Novelties

The talk contributes to the ongoing debate on whether LLMs can have Theory of Mind by proposing a framework of ‘differential construct validity’ to interpret LLM performance on human tasks. It also highlights the importance of using open-weight models for reproducibility and generalizability. The speaker’s perspective as a cognitive scientist using LLMs as model organisms offers a unique methodological angle.

Pour aller plus loin :

96 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical level, indicating a talk that is informative and credible but accessible to a broader audience.

Reliability 8/10

💬 No comments were provided.