Day 2: Health Sciences Grant Awardee Presentation #3 - Sophia Rein | ADIA Lab Symposium 2025

Day 2: Health Sciences Grant Awardee Presentation #3 - Sophia Rein | ADIA Lab Symposium 2025

🎙 Sophia Rein 👥 824 📅 November 5, 2025 ⏱ 10 min 👁 26 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

causal AIG-formulaLSTMreal-world evidencetarget trial emulation

Summary

Sophia Rein presents her ADIA Lab Health Sciences grant project on developing causal AI tools for generating real-world evidence from healthcare data. She begins by highlighting the abundance of observational health data (EHRs, claims, registries) and the gap in validated analytic tools to estimate causal effects at scale. The project aims to build an open-source software tool based on the G-formula, using deep learning (specifically LSTMs) to model complex longitudinal confounders, overcoming limitations of parametric models. The software will incorporate sample splitting and cross-fitting for valid inference and uncertainty quantification. The methodology involves splitting data, training LSTMs for confounders and outcome, estimating risks under sustained treatment strategies, and averaging cross-fitted estimates. The application focuses on comparing GLP-1 receptor agonists versus other anti-hypoglycemic drugs on cardiovascular events (MACE-4) in type 2 diabetes patients, using US (MarketScan) and UAE (Malaffi) data, emulating target trials. The project aims to provide accessible tools for health systems and foster US-UAE collaboration.

156 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation offers high value by addressing a critical need: translating vast healthcare data into actionable evidence. The argumentation is solid, grounded in established causal inference theory (G-formula) and modern machine learning (LSTMs). The proposed method is innovative in combining deep learning with causal inference for longitudinal data, potentially improving accuracy over parametric models. The application to a clinically relevant question (GLP-1 vs other drugs) is well-justified, given the limitations of RCTs. The inclusion of uncertainty quantification and simulation validation strengthens the scientific rigor. However, as a proposal, the actual impact is yet to be demonstrated.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is high: the methodology is clearly described, and the use of target trial emulation and cross-fitting aligns with best practices. The title accurately reflects the content. No external sources are cited in the video, but the description provides no links either. The presentation is a research proposal, so the lack of citations is acceptable, but it limits the ability to verify claims. The adequacy between title and content is perfect.

184 words

Title / Content Match

The title accurately reflects the content: a grant awardee presentation on health sciences, specifically causal AI for real-world evidence.

Quality & Reliability

8/10

The presentation is a research proposal by a Harvard-affiliated researcher, outlining a clear methodology (deep learning-based G-formula estimator) and a concrete application (GLP-1 vs other anti-hypoglycemic drugs on cardiovascular events). The approach is grounded in established causal inference theory (G-formula, NICE) and uses validated techniques (sample splitting, cross-fitting). However, it is a proposal, not yet validated results, and no external sources are cited in the video.

Key Moments

Contribution & Novelties

The project proposes a novel open-source software tool that integrates deep learning (LSTMs) with the G-formula for causal inference from longitudinal healthcare data, addressing the limitations of parametric models. This could significantly improve the accuracy of causal effect estimates in real-world settings. The application to GLP-1 vs other drugs is clinically relevant and leverages large-scale data from both the US and UAE.

Pour aller plus loin :

  • G-formula — Provides background on the causal inference framework used.
  • Long short-term memory (LSTM) — Explains the deep learning architecture employed.
  • Target trial emulation — Discusses the methodology for emulating randomized trials with observational data.

102 words

Radar Profile

The radar profile shows high scores in quality of information, technical level, and global reliability, reflecting the solid scientific foundation and clear methodology. The quantity of information is moderate, as it is a concise presentation of a research plan. Overall, the profile indicates a well-structured and credible proposal.

Reliability 8/10