「AIを応用した英語文献自動探索システムの開発とDNAバンク業務への実装」三輪 佳宏(理化学研究所 バイオリソース研究センター )

「AIを応用した英語文献自動探索システムの開発とDNAバンク業務への実装」三輪 佳宏(理化学研究所 バイオリソース研究センター )

🎙 三輪 佳宏 (Miwa Yoshihiro) 👥 123 📅 August 27, 2025 ⏱ 24 min 👁 48 📄 expert opinion 🧭 2026-08-17
Available in: English (current) Français

Keywords

AIliterature searchDNA bankdeep learningannotation

Summary

The presentation by Yoshihiro Miwa from RIKEN BRC describes the development and implementation of an AI-based system for automatically screening English scientific literature to support DNA bank operations. The system targets two types of papers: ‘resource development papers’ (describing new genetic resources) and ‘resource utilization papers’ (using provided resources). For the former, the team manually annotated 4000 papers from Scientific Reports over five years, defining four categories of genetic resource development. They used a PubMedBERT model with a 512-token limit, leading to the ‘subsection method’ where the Materials and Methods section is divided into subsections and each is scored, taking the highest score. This achieved high accuracy with false positive and false negative rates around 10%. The system can read 100 papers per minute, vastly outperforming human speed. Implementation increased overseas deposit requests from under 10 to about 100 per year, significantly boosting resource collection. For resource utilization papers, challenges included inconsistent author names and bank name variations, requiring improved traceability. The author also discusses issues with OpenAI’s filtering of certain terms (e.g., ‘sacrifice’) that may hide relevant papers, and notes that Elsevier’s ScienceDirect AI service is powerful but costly, with Japan lagging in AI adoption. The talk concludes with a call for researchers to use AI in literature exploration to avoid being left behind.

215 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation provides valuable insights into the practical application of AI for literature screening in a biological resource center. The author clearly explains the process from data annotation to model development and implementation, highlighting challenges such as token limits and the need for careful annotation guidelines. The argumentation is based on personal experience and specific examples, making it credible. However, the lack of quantitative performance metrics beyond the mentioned 10% error rates and the absence of comparison with other methods limit the depth of the argument. The discussion of industry practices (e.g., Elsevier’s pricing) is anecdotal but raises important considerations for researchers.

Scientific Rigor, Source Quality, Title Accuracy

The presentation is scientifically rigorous in its methodology, with a clear description of the annotation process and model development. The author does not cite specific external sources but refers to the use of PubMedBERT and mentions the ‘cat paper’ from Google, indicating familiarity with key AI developments. The title accurately reflects the content. The talk is based on the author’s own work, which adds credibility but also limits the generalizability. The mention of the NBRP workshop and the availability of a related video on YouTube provides additional context. Overall, the sources are appropriate for an expert opinion presentation, though more formal citations would enhance rigor.

222 words

Title / Content Match

The title accurately reflects the content, which details the development and implementation of an AI-based literature search system for DNA bank operations.

Quality & Reliability

7/10

The presentation is based on the author's direct experience developing an AI system for literature screening in a DNA bank context. It provides specific details on methodology, data annotation, and performance metrics, but lacks external validation or peer-reviewed publication of the system. The claims about AI limitations and industry practices are anecdotal.

Key Moments

Cited Sources

Concurring Sources

  • NBRP Workshop 2025 — The presentation was part of this workshop, providing context for the work.

Contribution & Novelties

The presentation offers a practical case study of applying AI to a specific biological resource management task, detailing the entire pipeline from data annotation to model deployment. It highlights the importance of domain-specific annotation and the challenges of token limits, leading to the innovative ‘subsection method’ for handling long documents. The discussion of AI filtering issues and the comparison of international AI adoption rates provides a broader perspective on the integration of AI in scientific research.

Pour aller plus loin :

  • PubMedBERT — A biomedical language model used in the system, relevant for understanding the underlying technology.
  • Deep learning — The general concept of deep learning, which is central to the AI development described.
  • Biobank — The context of DNA banks and biobanking, relevant to the application domain.
  • ScienceDirect AI — The AI-powered literature search service mentioned, useful for exploring current tools.

142 words

Radar Profile

The radar profile shows a balanced performance across all dimensions, with slightly higher scores in information quantity and quality, reflecting the detailed account of the AI system development. The technical level is moderate, suitable for a general scientific audience, while reliability is solid due to the first-hand experience shared.

Reliability 7/10