Lecture 9: Mini-Workshop: Automated Information Extraction from Electronic Medical Records, Aug 27, 11

Lecture 9: Mini-Workshop: Automated Information Extraction from Electronic Medical Records, Aug 27, 11

🎙 Dra. Helena Gómez Adorno 👥 4K 📅 August 28, 2025 ⏱ 64 min 👁 101 📄 tutorial 🧭 2026-08-13
Available in: English (current) Français

Keywords

clinical entity recognitionNLP pipelinerule-based systemsmachine learningLLMs

Summary

The lecture, part of a Mexico-Germany Hybrid Summer School on Medical Informatics with AI, focuses on automated information extraction from Spanish electronic medical records. The speaker, Dr. Helena Gómez Adorno, introduces the concept of information extraction as a subfield of NLP, transforming unstructured clinical text into structured data. She explains clinical entity recognition, identifying entities like diseases, medications, procedures, and vital signs. The talk covers the motivation for such extraction, including clinical decision support, epidemiological surveillance, and research. A traditional pipeline is presented: preprocessing (tokenization), named entity recognition, negation detection, relation extraction, and normalization to standard codes like ICD or SNOMED CT. Computational approaches are compared: rule-based systems, machine learning (e.g., BERT), and LLMs, highlighting trade-offs in interpretability, data requirements, privacy, and cost. The speaker shares a real-world application: a system developed for Mexico City’s health department (SEDESA) to extract vital signs, drugs, and comorbidities from COVID-19 clinical notes, using a combination of techniques. The talk concludes with a hands-on tutorial using a Google Colab notebook, demonstrating a rule-based NER system with spaCy and a Spanish model, and addressing challenges like negation detection.

183 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides valuable insights into the practical application of NLP in clinical settings, particularly for Spanish medical records. The speaker effectively argues for the use of machine learning over rule-based systems by highlighting the limitations of the latter in handling the variability and errors in clinical text, and the need for generalizable models. The argumentation is supported by a real-world case study, demonstrating the feasibility and benefits of automated extraction. The comparison of approaches (rule-based, ML, LLMs) is balanced, acknowledging trade-offs in interpretability, data requirements, privacy, and cost. The tutorial component adds practical value, allowing attendees to implement a basic NER system. However, the argumentation could be strengthened by more detailed evidence from the cited papers and a deeper discussion of evaluation metrics.

Scientific Rigor, Source Quality, Title Accuracy

The lecture demonstrates scientific rigor by presenting a clear methodology and referencing relevant work, including the speaker’s own publications. The sources cited are appropriate and include peer-reviewed papers and a practical notebook. The title accurately reflects the content, which is a mini-workshop on automated information extraction from electronic medical records. The presentation is well-structured, with a logical flow from concepts to applications and a hands-on exercise. The speaker acknowledges limitations and trade-offs, enhancing credibility. However, the lecture does not provide a systematic review of the literature, and some claims lack explicit citations. The practical examples are illustrative but not exhaustive. Overall, the scientific quality is high, and the title-content alignment is strong.

251 words

Title / Content Match

The title accurately reflects the content: a mini-workshop on automated information extraction from electronic medical records, with a focus on Spanish clinical notes.

Quality & Reliability

8/10

The lecture is given by a researcher with practical experience in the field, presenting a clear methodology and referencing real projects and publications. The content is well-structured and includes practical examples, but lacks detailed citations for all claims and does not provide a systematic review of the literature.

Key Moments

Cited Sources

  • Notebook for the tutorial — Referenced during the talk for the hands-on exercise
  • Paper on clinical entity recognition with domain adaptation — Mentioned as a recent publication by the speaker's lab
  • Paper on obstetric entity recognition for Robson criteria — Mentioned as a recent publication by the speaker's lab

Concurring Sources

Dissenting Sources

  • Potential biases in LLM-based extraction — The lecture acknowledges privacy and cost issues with LLMs but does not discuss potential biases in extraction, which is a known concern.

Contribution & Novelties

The lecture provides a comprehensive overview of automated information extraction from Spanish clinical notes, combining theoretical foundations with a practical tutorial. The speaker shares insights from a real-world implementation during the COVID-19 pandemic, highlighting the challenges and solutions in handling diverse clinical text. The tutorial offers a hands-on approach to building a rule-based NER system, which is valuable for beginners. The discussion of trade-offs between rule-based, ML, and LLM approaches is particularly useful for practitioners deciding on methodologies.

Pour aller plus loin :

  • Clinical NLP resources — Relevant for further reading on clinical natural language processing.
  • spaCy documentation — Official documentation for the library used in the tutorial.
  • SNOMED CT — Standard terminology for clinical entities, mentioned in the lecture.

120 words

Radar Profile

The radar profile shows high scores in quantity and quality of information, and moderate technical level, indicating a well-balanced lecture that is informative and practical. The fiabilité is high, reflecting the speaker's expertise and real-world experience.

Reliability 8/10

💬 Sur les 0 commentaires analysés, aucune tendance n'est disponible.