LLMs in Science: The Good, The Bad and The Ugly - Nihar Shah

LLMs in Science: The Good, The Bad and The Ugly - Nihar Shah

🎙 Nihar Shah 👥 4K 📅 December 9, 2025 ⏱ 66 min 👁 61 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

LLMpeer reviewfraudAI scientistreviewer assignmentbiasevaluation

Summary

Nihar Shah, an associate professor at Carnegie Mellon University, presents a seminar on the integration of large language models (LLMs) into scientific processes, focusing on peer review. He structures his talk into three parts: ‘The Good’, ‘The Bad’, and ‘The Ugly’. In ‘The Good’, he discusses the potential of LLMs to identify errors in papers, presenting experiments where human reviewers often miss fundamental flaws, while LLMs can catch some with appropriate prompting. He also highlights biases in human evaluation of reviews, such as authors favoring positive reviews and reviewers being swayed by verbosity. In ‘The Bad’, he exposes vulnerabilities in the peer review system to fraud, including fake accounts and collusion rings, and demonstrates how text-matching-based reviewer assignment can be manipulated. In ‘The Ugly’, he addresses methodological pitfalls in autonomous ‘AI scientists’, emphasizing the need for rigorous evaluation. The talk is based on empirical studies and randomized controlled trials, providing evidence-based insights into the challenges and opportunities of using LLMs in science.

162 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the application of LLMs in peer review, supported by empirical data from controlled experiments. Shah presents concrete examples and statistics, such as the low detection rate of errors by human reviewers (only 1 out of 79 identified a subtle error) and the bias in author evaluations. The argumentation is solid, with clear explanations of experimental designs and limitations. He also discusses the broader implications for scientific integrity, making the content highly relevant for researchers and policymakers.

Scientific Rigor, Source Quality, Title Accuracy

Shah demonstrates scientific rigor by referencing published studies and ongoing research, including experiments at NeurIPS 2022 and collaborations with program chairs. He acknowledges limitations, such as the generalizability of single-paper experiments. The title accurately reflects the content, and the talk is well-structured. The sources cited are credible, and the speaker’s expertise adds to the reliability. The adéquation between title and content is strong, with no significant discrepancies.

164 words

Title / Content Match

The title accurately reflects the three-part structure of the talk, covering positive, negative, and problematic aspects of LLMs in science.

Quality & Reliability

8/10

The talk presents empirical studies and randomized controlled trials from peer-reviewed venues, with detailed methodology and transparent limitations. The speaker is a recognized expert in the field, and the content is grounded in published research, though some claims are based on unpublished or ongoing work.

Key Moments

Cited Sources

  • OpenReview — Mentioned as a platform with publicly available reviews and as a venue where fake profiles were found.
  • NeurIPS 2022 — The venue where the randomized controlled trial on review evaluation was conducted.
  • ICLR — Mentioned as a conference with publicly available past data, used for training attacks.
  • CVPR — Mentioned as a venue that uses only text matching for reviewer assignment.
  • ARR (Association for Computational Linguistics Rolling Review) — Mentioned as a venue that has banned bidding and uses only text matching.

Concurring Sources

  • The Science of Peer Review — An article discussing the challenges and potential improvements in peer review, aligning with the talk's themes.
  • AI for peer review — An article exploring the use of AI in peer review, supporting the potential benefits and risks discussed.

Dissenting Sources

Contribution & Novelties

The talk provides novel empirical evidence on the limitations of human peer review and the potential of LLMs to assist in error detection. It also highlights vulnerabilities in the peer review system to fraud, which is an underexplored area. The speaker’s approach of using controlled experiments to evaluate AI reviewers is innovative and contributes to the field of science of science.

Pour aller plus loin :

134 words

Radar Profile

The radar profile shows high scores in quality of information and reliability, reflecting the speaker's expertise and the empirical nature of the talk. The quantity of information is also high, but the technical level is moderate, making it accessible to a broad audience. The overall balance indicates a well-rounded presentation with strong scientific grounding.

Reliability 8/10

💬 Sur les 0 commentaires analysés, aucune tendance n'a pu être dégagée.