The munge function

The munge function

🎙 International Statistical Genetics Workshop 👥 3K 📅 May 18, 2026 ⏱ 16 min 👁 253 📄 tutorial 🧭 2026-08-16
Available in: English (current) Français

Keywords

mungegenomic SEMGWASLD score regressiondata preprocessing

Summary

This video is a technical tutorial on the munge function, the first step in the genomic SEM pipeline. It explains the importance of data munging for preparing GWAS summary statistics for LD score regression. The presenter outlines the five key pieces of information required in GWAS data: RSID, effect allele (A1), non-effect allele (A2), a signed effect column (e.g., beta or odds ratio), and p-value. He emphasizes the need for ancestry matching across datasets and LD scores, and provides guidelines for power, such as including traits with SNP-based heritability Z > 4. The video details how to handle sample size, particularly for binary traits, using the liability scale correction and effective sample size. It demonstrates how to back out effective sample size from data when not directly available, using an equation from a referenced publication. The munge function’s arguments are explained, including the hapmap3 reference file for allele alignment. The presenter stresses the importance of inspecting the output log file for quality control. The tutorial concludes with a practical R code example, showing how to run munge on a major depression GWAS dataset.

183 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides high-value information for researchers using genomic SEM, offering practical guidance on data preparation that is often overlooked. The argumentation is solid, based on established statistical genetics principles and referencing relevant literature. The presenter clearly explains complex concepts like liability scale correction and effective sample size, making them accessible. The use of a real example (MDD) and code demonstration enhances the practical value. The video also addresses common pitfalls, such as ancestry matching and sample size calculation, which are crucial for valid results.

94 words

Title / Content Match

The title accurately reflects the content, which focuses on the munge function in the genomic SEM pipeline.

Quality & Reliability

8/10

The video is a technical tutorial by an expert in statistical genetics, providing detailed and accurate information about data munging for genomic SEM. It references specific methods and publications, and includes practical code examples. The content is well-structured and aligns with established practices in the field.

Key Moments

Cited Sources

  • Genomic SEM GitHub — Referenced as a resource for data sources and package documentation.
  • Biological Psychiatry publication on effective sample size — Cited for the equation to back out effective sample size.

Concurring Sources

  • Genomic SEM paper — Provides the theoretical foundation for genomic SEM, which the video builds upon.
  • LDSC documentation — Official LDSC GitHub, which includes guidelines on data munging and effective sample size.

Contribution & Novelties

This video provides a detailed, step-by-step guide to the munge function, which is often a bottleneck in genomic SEM analyses. It clarifies common pitfalls and offers practical solutions, such as backing out effective sample size when not directly available. The tutorial is valuable for researchers new to genomic SEM, as it demystifies the data preparation process.

Pour aller plus loin :

84 words

Radar Profile

The radar profile shows balanced scores across all dimensions, indicating a well-rounded tutorial with strong technical depth and reliability. The high scores in information quantity and quality reflect the comprehensive coverage of the munge function, while the technical level is appropriate for the target audience.

Reliability 8/10