CAOS 2025 - 6 | Rovereto, May 7-9 | Jeffrey S. Bowers

CAOS 2025 - 6 | Rovereto, May 7-9 | Jeffrey S. Bowers

Humanities, Social Sciences & Thought Psychology JMPsychologyJMRCognition and cognitive psychology
🎙 Jeffrey S. Bowers 👥 2K 📅 November 15, 2025 ⏱ 71 min 👁 31 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

neural networksvisionlanguagebrain scoremechanistic claims

Summary

Jeffrey S. Bowers presents a critical analysis of the claims that artificial neural networks (ANNs) provide mechanistic insights into human vision and language. He argues that the field often relies on correlational studies (e.g., brain score benchmarks) that do not support causal or mechanistic conclusions. He highlights that brain score measures are often driven by confounds, such as background information rather than object identity, and that behavioral benchmarks can be misleading when models are trained on large datasets. He also criticizes the tendency to make strong claims from weak findings, citing examples where models explain only a tiny fraction of variance in brain responses. Bowers emphasizes the importance of running controlled experiments to test specific hypotheses, as is standard in psychology, and suggests that the current culture of the field favors publishing claims of similarity over differences, leading to an overestimation of the relevance of ANNs to understanding human cognition.

150 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides a valuable critical perspective on the use of neural networks in cognitive neuroscience. Bowers effectively argues that correlational approaches, such as brain score, are insufficient for establishing mechanistic similarity. He supports his claims with concrete examples, including his own research showing that brain score predictions are largely driven by background confounds, and he references studies where models fail on out-of-distribution stimuli that humans find easy. The argumentation is coherent and well-structured, though it is primarily an opinion piece rather than a systematic review. The speaker’s expertise in the field adds credibility, and he acknowledges the importance of experiments in psychology. However, the talk could benefit from a more balanced discussion of successful model-brain alignments, as it focuses heavily on failures.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor in its use of published studies and the speaker’s own research. Bowers cites specific papers, such as the PNAS study on transformer models and the Vong et al. Science paper on language acquisition, and he discusses the limitations of these studies. The quality of sources is high, and the arguments are based on empirical evidence. The title accurately reflects the content, which is a critical examination of neural network models. The talk does not include a discussion of public comments, as none were provided.

227 words

Title / Content Match

The title accurately reflects the content, which focuses on deep problems with neural network models of human vision and language.

Quality & Reliability

8/10

The talk presents a critical perspective on neural network models of vision and language, supported by references to published studies and the speaker's own research. The arguments are logically structured and based on empirical evidence, though the presentation is an opinion piece rather than a systematic review.

Key Moments

Cited Sources

Concurring Sources

Dissenting Sources

Contribution & Novelties

The talk offers a critical perspective on the use of neural networks in cognitive neuroscience, highlighting methodological flaws in common approaches like brain score. It emphasizes the importance of controlled experiments over correlational studies. The speaker provides original examples from his own research demonstrating confounds in brain score benchmarks.

Pour aller plus loin :

  • Brain-Score — A benchmark platform for comparing models to brain data, central to the talk’s critique.
  • Texture vs. Shape Bias in CNNs — Paper by Geirhos et al. showing that CNNs rely on texture, a key example in the talk.
  • Vong et al. (2024) on grounded language acquisition — Study on language acquisition from a single child’s perspective, discussed in the talk.

116 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical level. This indicates a well-supported critical talk that is accessible to a broad scientific audience.

Reliability 8/10