Polygenic Prediction: Part 2 Evaluation, visualization, and pitfalls of polygenic prediction

Polygenic Prediction: Part 2 Evaluation, visualization, and pitfalls of polygenic prediction

🎙 International Statistical Genetics Workshop 👥 3K 📅 May 18, 2026 ⏱ 21 min 👁 354 📄 tutorial 🧭 2026-08-16
Available in: English (current) Français

Keywords

polygenic risk scoreR-squaredAUCliability threshold modeloverfitting

Summary

This lecture, part of a series on polygenic prediction, focuses on evaluating the performance of polygenic scores (PGS). It begins by discussing metrics for quantitative traits, primarily the prediction R-squared, which measures the squared correlation between observed phenotype and PGS in an independent sample. The lecture emphasizes that prediction R-squared is bounded by SNP-based heritability and that incremental R-squared should be reported after adjusting for covariates. Visualization of predictions can help identify outliers and assess bias, with the expected slope of phenotype on PGS being 1 for unbiased predictions. For binary traits, the lecture introduces several metrics: pseudo R-squared from logistic regression, AUC, and variance explained on the liability scale. Pseudo R-squared is criticized for its dependence on sample prevalence, while AUC is independent of case-control proportion but its maximum value depends on disease prevalence. The liability scale R-squared, based on the liability threshold model, is recommended as it is independent of prevalence and comparable to heritability estimates. The lecture also covers risk stratification and decile odds ratios for visualization, noting their sample dependence. Finally, it highlights common pitfalls, including in-sample prediction, sample overlap, and SNP selection using the target sample, which can inflate accuracy. The importance of independent discovery and target samples is stressed, with cross-dataset prediction as the gold standard.

212 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides a comprehensive and practical overview of evaluation metrics for polygenic prediction, with clear explanations of each metric’s properties and limitations. The argumentation is solid, supported by theoretical explanations and graphical illustrations. The discussion of pitfalls is particularly valuable, as it addresses common errors that can lead to inflated results, and the recommendation of liability-scale R-squared is well-justified. The lecture effectively balances theoretical foundations with practical considerations, making it a useful resource for researchers in statistical genetics.

Scientific Rigor, Source Quality, Title Accuracy

The lecture demonstrates scientific rigor by referencing key literature, such as Lee et al. (2012) for the liability threshold model, and by explaining the theoretical basis for each metric. The quality of sources is adequate, though the lecture does not provide a comprehensive reference list. The title accurately reflects the content, which focuses on evaluation, visualization, and pitfalls. The presentation is clear and well-structured, with visual aids that enhance understanding. Overall, the scientific rigor is high, and the content aligns with the title.

177 words

Title / Content Match

The title accurately reflects the content, which focuses on evaluation metrics, visualization, and pitfalls in polygenic prediction.

Quality & Reliability

8/10

The lecture is well-structured, covers standard statistical genetics methods, and highlights common pitfalls. It references key literature (Lee et al. 2012) and provides practical guidance. However, it lacks explicit citations for some claims and does not provide detailed derivations.

Key Moments

Cited Sources

  • Lee et al. 2012 - Liability threshold model — Referenced for the theory of mapping observed scale R-squared to liability scale.

Concurring Sources

  • Lee et al. 2012 - Liability threshold model — The lecture's explanation of liability scale R-squared aligns with the theoretical framework presented in this paper.

Contribution & Novelties

This lecture provides a clear and structured overview of evaluation metrics for polygenic prediction, with a strong emphasis on practical pitfalls. It uniquely highlights the importance of liability-scale R-squared as the most interpretable metric, and provides concrete examples of how sample ascertainment can affect results. The discussion of pitfalls, such as in-sample prediction and sample overlap, is particularly valuable for researchers.

Pour aller plus loin :

97 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and reliable educational resource. The lecture excels in providing quantitative information and technical depth, with a strong focus on methodological rigor.

Reliability 8/10