Keywords
Summary
215 words
Critical Evaluation
Value of the Information & Strength of the Argument
The presentation provides valuable insights into the practical application of AI for literature screening in a biological resource center. The author clearly explains the process from data annotation to model development and implementation, highlighting challenges such as token limits and the need for careful annotation guidelines. The argumentation is based on personal experience and specific examples, making it credible. However, the lack of quantitative performance metrics beyond the mentioned 10% error rates and the absence of comparison with other methods limit the depth of the argument. The discussion of industry practices (e.g., Elsevier’s pricing) is anecdotal but raises important considerations for researchers.
Scientific Rigor, Source Quality, Title Accuracy
The presentation is scientifically rigorous in its methodology, with a clear description of the annotation process and model development. The author does not cite specific external sources but refers to the use of PubMedBERT and mentions the ‘cat paper’ from Google, indicating familiarity with key AI developments. The title accurately reflects the content. The talk is based on the author’s own work, which adds credibility but also limits the generalizability. The mention of the NBRP workshop and the availability of a related video on YouTube provides additional context. Overall, the sources are appropriate for an expert opinion presentation, though more formal citations would enhance rigor.
222 words
Title / Content Match
The title accurately reflects the content, which details the development and implementation of an AI-based literature search system for DNA bank operations.
Quality & Reliability
7/10
The presentation is based on the author's direct experience developing an AI system for literature screening in a DNA bank context. It provides specific details on methodology, data annotation, and performance metrics, but lacks external validation or peer-reviewed publication of the system. The claims about AI limitations and industry practices are anecdotal.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction: Speaker introduces himself and the topic of AI-based literature search system for DNA bank.
- Explanation of AI development basics: history of AI winters and the 2012 breakthrough with deep learning (Google cat paper).
- Description of the supervised learning approach: annotation of 4000 papers from Scientific Reports by 7 team members.
- Details on the definition of genetic resources and the four categories used for annotation.
- Discussion of token limit (512 tokens) and the need to split papers into subsections (subsection method).
- Performance results: false positive and false negative rates around 10%, and speed of 100 papers per minute.
- Implementation in DNA bank operations: increase in overseas deposit requests from <10 to ~100 per year.
- Challenges with resource utilization papers: inconsistent author names and bank name variations, need for traceability.
- Issues with OpenAI filtering: suspicion that certain terms (e.g., 'sacrifice') are filtered, hiding relevant papers.
- Discussion of Elsevier's ScienceDirect AI service and the high cost of data access; Japan's lag in AI adoption.
- Conclusion: Emphasis on the necessity of using AI for literature exploration to avoid being left behind.
Cited Sources
- NBRP Workshop 2025 — The presentation was part of this workshop.
Concurring Sources
- NBRP Workshop 2025 — The presentation was part of this workshop, providing context for the work.
Contribution & Novelties
The presentation offers a practical case study of applying AI to a specific biological resource management task, detailing the entire pipeline from data annotation to model deployment. It highlights the importance of domain-specific annotation and the challenges of token limits, leading to the innovative ‘subsection method’ for handling long documents. The discussion of AI filtering issues and the comparison of international AI adoption rates provides a broader perspective on the integration of AI in scientific research.
Pour aller plus loin :
- PubMedBERT — A biomedical language model used in the system, relevant for understanding the underlying technology.
- Deep learning — The general concept of deep learning, which is central to the AI development described.
- Biobank — The context of DNA banks and biobanking, relevant to the application domain.
- ScienceDirect AI — The AI-powered literature search service mentioned, useful for exploring current tools.
142 words
Radar Profile
The radar profile shows a balanced performance across all dimensions, with slightly higher scores in information quantity and quality, reflecting the detailed account of the AI system development. The technical level is moderate, suitable for a general scientific audience, while reliability is solid due to the first-hand experience shared.
