Artificial Cognition in LLMs

Artificial Cognition in LLMs

🎙 Karin de Langis 👥 843 📅 August 12, 2026 ⏱ 60 min 👁 10 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

artificial cognitionLLMworking memoryexecutive functioncognitive science

Summary

Karin de Langis presents her research on artificial cognition in large language models (LLMs), applying methods from cognitive science to understand whether LLMs exhibit genuine reasoning or sophisticated pattern matching. She introduces the concept of the ‘jagged frontier’ of AI capabilities, highlighting unpredictable failures and the need for better understanding. The talk focuses on executive functions, particularly working memory, and describes a series of experiments adapted from human psychology paradigms, including operation span, reading span, digit span, and n-back tasks. Results show that LLMs generally outperform humans in working memory tasks, except for n-back where performance is similar. However, LLMs show a significant cost in reversing in-context information (backward digit span). The speaker also mentions the Wisconsin Card Sorting Test for cognitive flexibility, but does not detail results due to time. The overarching goal is to identify bottlenecks in LLM problem-solving and to use cognitive science as a framework for understanding AI capabilities and limitations.

155 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the cognitive capabilities of LLMs, using established experimental paradigms from cognitive psychology. The argumentation is solid, grounded in prior literature and careful experimental design. The speaker acknowledges the preliminary nature of some findings and discusses potential implications for understanding AI failures. The comparison with human baselines is informative, and the discussion of the ‘jagged frontier’ provides a compelling motivation for the research. The presentation is well-structured and the speaker encourages questions, fostering engagement.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor by referencing prior work (e.g., Gong et al. 2024) and using established cognitive tasks. The speaker is transparent about the methodology, including prompt variations and data formatting. The title accurately reflects the content. The sources cited are primarily academic papers mentioned in the talk, though specific URLs are not provided in the description. The talk is suitable for a technical audience, but the speaker does not oversimplify the concepts.

168 words

Title / Content Match

The title accurately reflects the content, which focuses on characterizing artificial cognition in LLMs through cognitive science paradigms.

Quality & Reliability

8/10

The talk presents original research from a PhD candidate at a reputable university, with methods grounded in cognitive science and results compared to established human baselines. The speaker is transparent about limitations and the preliminary nature of some findings. The presentation is clear and well-structured, though it lacks peer-reviewed publication details for some results.

Key Moments

Cited Sources

  • Gong et al. (2024) - n-back task on GPT-4 — Referenced as prior work on n-back performance in LLMs.
  • Jagged Frontier paper (2023) — Coined the term 'jagged frontier' for AI capabilities.

Concurring Sources

  • Gong et al. (2024) - n-back task on GPT-4 — Replicated in this talk with multiple LLMs.

Contribution & Novelties

The talk contributes original empirical results on LLM working memory and executive functions, replicating and extending prior work. It provides a systematic comparison across multiple LLMs and tasks, highlighting strengths and weaknesses. The use of cognitive science paradigms offers a novel lens for understanding AI behavior.

Pour aller plus loin :

90 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical level, indicating a well-balanced presentation that is accessible yet rigorous.

Reliability 8/10

💬 No comments were provided for analysis.