
LLMs in Science: The Good, The Bad and The Ugly - Nihar Shah
Keywords
Summary
162 words
Critical Evaluation
Value of the Information & Strength of the Argument
The talk provides valuable insights into the application of LLMs in peer review, supported by empirical data from controlled experiments. Shah presents concrete examples and statistics, such as the low detection rate of errors by human reviewers (only 1 out of 79 identified a subtle error) and the bias in author evaluations. The argumentation is solid, with clear explanations of experimental designs and limitations. He also discusses the broader implications for scientific integrity, making the content highly relevant for researchers and policymakers.
Scientific Rigor, Source Quality, Title Accuracy
Shah demonstrates scientific rigor by referencing published studies and ongoing research, including experiments at NeurIPS 2022 and collaborations with program chairs. He acknowledges limitations, such as the generalizability of single-paper experiments. The title accurately reflects the content, and the talk is well-structured. The sources cited are credible, and the speaker’s expertise adds to the reliability. The adéquation between title and content is strong, with no significant discrepancies.
164 words
Title / Content Match
The title accurately reflects the three-part structure of the talk, covering positive, negative, and problematic aspects of LLMs in science.
Quality & Reliability
8/10
The talk presents empirical studies and randomized controlled trials from peer-reviewed venues, with detailed methodology and transparent limitations. The speaker is a recognized expert in the field, and the content is grounded in published research, though some claims are based on unpublished or ongoing work.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction and overview of the talk's structure: The Good, The Bad, The Ugly.
- Discussion on evaluating AI reviewers and the challenge of matching human reviews.
- Experiment with 79 reviewers: only 1 identified a subtle error, highlighting human review limitations.
- Analysis of review quality and the lack of correlation with self-reported confidence.
- Experiment at NeurIPS 2022 showing authors' bias in evaluating reviews.
- Randomized controlled trial demonstrating that longer reviews receive higher scores, indicating style bias.
- Comparison of LLM and human error detection: GPT-4 found obvious errors but missed subtle ones without steering.
- Introduction to fraud in peer review: fake accounts and identity theft.
- Collusion rings and the vulnerability of reviewer assignment systems.
- Attack vectors on text-matching-based assignment: modifying abstracts, updating profiles, and using past data.
Cited Sources
- OpenReview — Mentioned as a platform with publicly available reviews and as a venue where fake profiles were found.
- NeurIPS 2022 — The venue where the randomized controlled trial on review evaluation was conducted.
- ICLR — Mentioned as a conference with publicly available past data, used for training attacks.
- CVPR — Mentioned as a venue that uses only text matching for reviewer assignment.
- ARR (Association for Computational Linguistics Rolling Review) — Mentioned as a venue that has banned bidding and uses only text matching.
Concurring Sources
- The Science of Peer Review — An article discussing the challenges and potential improvements in peer review, aligning with the talk's themes.
- AI for peer review — An article exploring the use of AI in peer review, supporting the potential benefits and risks discussed.
Dissenting Sources
Contribution & Novelties
The talk provides novel empirical evidence on the limitations of human peer review and the potential of LLMs to assist in error detection. It also highlights vulnerabilities in the peer review system to fraud, which is an underexplored area. The speaker’s approach of using controlled experiments to evaluate AI reviewers is innovative and contributes to the field of science of science.
Pour aller plus loin :
- Peer review in scientific publications — Background on the peer review process.
- Large language models — Overview of LLMs and their capabilities.
- Scientific fraud — Discussion on types of scientific fraud, including collusion and identity theft.
- Specter model — The model used for text matching in reviewer assignment, as mentioned in the talk.
- Simpson’s paradox — A statistical phenomenon mentioned as an example of conceptual errors in papers.
134 words
Radar Profile
The radar profile shows high scores in quality of information and reliability, reflecting the speaker's expertise and the empirical nature of the talk. The quantity of information is also high, but the technical level is moderate, making it accessible to a broad audience. The overall balance indicates a well-rounded presentation with strong scientific grounding.
💬 Sur les 0 commentaires analysés, aucune tendance n'a pu être dégagée.