BioML Seminar 2.4 - Tianyu Lu on Evaluating Protein Design Models with SHAPES

BioML Seminar 2.4 - Tianyu Lu on Evaluating Protein Design Models with SHAPES

🎙 Tianyu Lu 👥 14K 📅 October 30, 2025 ⏱ 81 min 👁 305 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

SHAPESprotein designgenerative modelsdesignabilityFréchet Protein Distance

Summary

Tianyu Lu presents SHAPES (Structural and Hierarchical Assessment of Proteins with Embedding Similarity), a framework for evaluating generative models of protein structures. The motivation is that current models are biased towards designable, idealized structures, often oversampling alpha-helices and undersampling loops and flexible regions crucial for function. The study benchmarks five state-of-the-art generative models against the CATH database and the PDB, using structural embeddings across multiple hierarchies (local, intermediate, global) to compare sample distributions. They introduce the Fréchet Protein Distance (FPD) to quantify distributional coverage. Key findings include a systematic bias towards secondary structure elements, with RFdiffusion showing extreme helical bias. They also demonstrate that the designability metric, based on self-consistency, fails for 40-60% of native proteins, indicating a mismatch between computational proxy and experimental reality. The analysis reveals undersampling of complex, loop-rich structures and enzymes, and oversampling of idealized helical bundles. The talk includes a discussion of limitations, such as the Gaussian assumption in FPD, and potential implications for functional protein design.

162 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the limitations of current protein design evaluation methods. The argumentation is solid, supported by quantitative analyses and visualizations. The speaker clearly explains the designability metric and its pitfalls, and introduces a more nuanced evaluation framework (SHAPES) that captures structural diversity. The use of multiple embedding hierarchies and the FPD metric is well-justified. The discussion of undersampling and oversampling regions in structure space is compelling and backed by examples. The acknowledgment of limitations, such as the Gaussian assumption, adds to the credibility. Overall, the talk offers a significant contribution to the field by highlighting the need for more comprehensive evaluation metrics.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor through a clear methodology, careful selection of reference datasets, and acknowledgment of potential artifacts. The speaker cites relevant prior work, such as EvoDiff and the use of FID in image generation, and provides a preprint reference. The title accurately reflects the content. The talk does not include a promotional segment. The speaker is transparent about the limitations of the study, such as the Gaussian assumption in FPD and the backbone-only nature of the embeddings. The analysis of native proteins failing designability is a critical point that challenges current practices.

215 words

Title / Content Match

The title accurately reflects the content: a seminar talk by Tianyu Lu on evaluating protein design models using the SHAPES framework.

Quality & Reliability

8/10

The talk presents a rigorous benchmarking study with clear methodology, quantitative metrics, and critical analysis of existing evaluation practices. The speaker is a PhD candidate at Stanford, and the work is available as a preprint. Limitations are acknowledged, and the approach is well-motivated.

Key Moments

Cited Sources

Concurring Sources

  • EvoDiff paper — Used protein language model embeddings to compare generated and real sequences, inspiring the approach in SHAPES.

Contribution & Novelties

The talk introduces SHAPES, a novel evaluation framework for generative protein models that goes beyond designability by assessing distributional coverage across structural hierarchies. It provides a quantitative metric (FPD) and reveals systematic biases in current models. The finding that a significant fraction of native proteins fail the designability test challenges the field’s reliance on this metric.

Pour aller plus loin :

90 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-balanced and rigorous presentation. The talk excels in providing substantial information, technical depth, and reliable methodology, with a slight emphasis on quantitative analysis.

Reliability 8/10