BioML Seminar 3.2 - Sarah Gurev on Benchmarking Model Performance on Pandemic-Threat Viruses

BioML Seminar 3.2 - Sarah Gurev on Benchmarking Model Performance on Pandemic-Threat Viruses

🎙 Sarah Gurev 👥 14K 📅 October 30, 2025 ⏱ 61 min 👁 230 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

EVERESTprotein language modelsviral mutationbenchmarkdeep mutational scanning

Summary

Sarah Gurev presents a framework called EVEREST for benchmarking model performance on viral mutation effect prediction. She begins by motivating the need for early and accurate predictions of viral mutations, using the example of SARS-CoV-2 and the potential of using pre-2020 coronavirus sequences. She explains alignment-based models, including site-independent, pairwise, and variational autoencoder approaches, and demonstrates their utility in predicting antibody escape and vaccine effectiveness. She then introduces the EVEREST benchmark, which consists of 45 curated viral deep mutational scanning datasets, and uses it to evaluate both alignment-based models and protein language models. Key findings include that deeper alignments are not always better, and that protein language models fail to reliably predict mutations in over half of 40 WHO-prioritized pandemic-threat viruses. The talk concludes with actionable recommendations for improving viral mutation effect prediction and an objective framework for analyzing dual-use biosecurity risk.

142 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation provides valuable insights into the strengths and limitations of current computational models for viral mutation prediction. The argumentation is solid, supported by quantitative results from the EVEREST benchmark. The speaker carefully explains the methodology and highlights important caveats, such as the role of epistasis and the need for alignment relevance. The comparison between alignment-based models and protein language models is particularly informative, revealing that protein language models, despite their success in other domains, underperform on many viral families. The recommendations for improving model performance are actionable and grounded in the data.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is high, with a clear experimental design and use of established datasets. The speaker cites relevant prior work, including deep mutational scanning studies and protein language models like ESM and ProGen. The title accurately reflects the content. The talk is based on original research, and the methodology is transparent. The speaker acknowledges limitations, such as the difficulty of predicting all mutations and the need for experimental validation. The sources cited are appropriate and credible.

185 words

Title / Content Match

The title accurately reflects the content, which focuses on benchmarking model performance on pandemic-threat viruses.

Quality & Reliability

8/10

Presentation of original research with a clear methodology, benchmark construction, and quantitative results. The speaker is a domain expert with a strong academic background. Limitations and uncertainties are acknowledged, and the work is grounded in established datasets and models.

Key Moments

Cited Sources

  • EVEscape — Mentioned as a model for predicting antibody escape using evolutionary sequences.
  • Deep mutational scanning datasets — Used to benchmark model performance in EVEREST.
  • Protein language models (ESM, ProGen) — Discussed as alternative models for mutation effect prediction.

Concurring Sources

  • EVE model — Supports the use of evolutionary sequence models for variant effect prediction.

Contribution & Novelties

The talk introduces EVEREST, a comprehensive benchmark for evaluating viral mutation effect prediction models, which is a significant contribution to the field. It provides a systematic comparison of alignment-based and protein language models across diverse viral families, revealing important limitations of current approaches. The findings offer actionable recommendations for improving model performance and highlight the need for alignment relevance. The framework also addresses dual-use biosecurity risk, adding a critical dimension to the evaluation.

Pour aller plus loin :

  • Protein language models — Overview of language models applied to proteins.
  • Deep mutational scanning — Experimental technique used to measure mutation effects.
  • EVE model — Prior work on evolutionary model of variant effects.
  • AlphaFold — Structure prediction using co-evolutionary signals.

118 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a technically rigorous and well-sourced presentation. The balance between information quantity, quality, and technical depth suggests a comprehensive and reliable resource for experts.

Reliability 8/10

💬 Sur les 0 commentaires analysés, aucune tendance n'est disponible.