Behavioral evaluation of language models as models of human sentence processing

Behavioral evaluation of language models as models of human sentence processing

🎙 Roger Levy 👥 305 📅 September 18, 2025 ⏱ 85 min 👁 131 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

surprisallanguage modelsreading timesBayesian inferencecognitive modeling

Summary

Roger Levy, professor at MIT, presents a seminar on using large language models (LLMs) as computational models of human sentence processing. He introduces the concept of surprisal, an information-theoretic measure of word predictability, and shows how it correlates with human reading times. He discusses the importance of using direct probability measurements from LLMs rather than prompting-based evaluations, citing recent work. He also covers the rational analysis framework, Bayesian inference, and the noisy channel model. He presents evidence that LLMs can capture syntactic structures and predict human processing difficulty, but also acknowledges limitations. The talk includes examples of LLMs handling complex sentences and discusses the implications for cognitive science and NLP.

110 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the intersection of LLMs and psycholinguistics, offering a clear theoretical framework (rational analysis) and empirical evidence. The argumentation is solid, grounded in published research and logical reasoning. Levy effectively demonstrates how surprisal from LLMs aligns with human reading times, supporting the validity of these models as cognitive models. He also critically examines limitations, such as the need for direct probability measures over prompting. The presentation is well-structured, moving from theoretical foundations to concrete examples and recent findings.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, with references to peer-reviewed publications (e.g., Hu & Levy 2023, Shain et al. 2024, Wilcox et al. 2023, Futrell et al. 2020). The sources are credible and directly support the claims. The title accurately reflects the content, focusing on behavioral evaluation of language models. The presentation is well-organized and the speaker’s expertise is evident. No comments were provided, so no analysis of public reception is included.

169 words

Title / Content Match

The title accurately reflects the content, which focuses on behavioral evaluation methods for language models in the context of human sentence processing.

Quality & Reliability

9/10

Presentation by a leading expert (MIT professor, president of Cognitive Science Society) with references to peer-reviewed publications. The talk is a synthesis of established and recent research, with clear methodological explanations.

Key Moments

Cited Sources

  • Prompting is not a substitute for probability measurements in large language models — Cited as evidence that direct probability measurements are better than prompting for evaluating LLMs.
  • Large-scale evidence for logarithmic effects of word predictability on reading time — Cited as evidence for the relationship between surprisal and reading times.
  • Using Computational Models to Test Syntactic Learnability — Cited as work using computational models to test syntactic knowledge.
  • Lossy-context surprisal: An information-theoretic model of memory effects in sentence processing — Cited as a model of memory effects in sentence processing.

Concurring Sources

Contribution & Novelties

The talk provides a comprehensive overview of recent advances in using LLMs as cognitive models, emphasizing the importance of behavioral evaluation methods. It highlights the superiority of direct probability measurements over prompting for assessing linguistic knowledge. The presentation of surprisal as a key metric linking LLM predictions to human processing difficulty is a significant contribution. The talk also discusses the rational analysis framework and its application to language processing, offering a theoretical grounding for LLM-based models.

Pour aller plus loin :

  • Surprisal — Wikipedia article on surprisal, a key concept in information theory and psycholinguistics.
  • Bayesian inference — Wikipedia article on Bayesian inference, central to the rational analysis framework.
  • Noisy channel model — Wikipedia article on noisy channel models, relevant to language processing under noise.

125 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-balanced and rigorous presentation. The talk excels in information quantity and quality, with strong technical depth and high reliability. The overall assessment is excellent.

Reliability 9/10