Johann D. Gaebler: A Simple, Statistically Robust Test of Discrimination

Johann D. Gaebler: A Simple, Statistically Robust Test of Discrimination

🎙 Johann D. Gaebler 👥 2K 📅 May 12, 2026 ⏱ 21 min 👁 145 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

discriminationoutcome testbenchmark testmonotone likelihood ratioinframarginality

Summary

Johann D. Gaebler presents a new statistical test for detecting discriminatory double standards, developed with Sharad Goel and published in PNAS. The talk begins by motivating the problem with a hypothetical bank lending scenario, highlighting the limitations of the traditional benchmark test (comparing approval rates) and Becker’s outcome test (comparing repayment rates). Both can yield misleading results due to omitted variables and infra-marginality. The proposed ‘robust outcome test’ combines these two tests: it concludes discrimination only when both the benchmark and outcome tests point in the same direction. The speaker formalizes this intuition using the monotone likelihood ratio property, showing that under this condition, the test is reliable. Empirical evidence from lending, criminal justice, and admissions data suggests that monotonicity often holds approximately. Simulation studies demonstrate that the robust outcome test is almost always correct, while the standard outcome test is frequently wrong. The method is illustrated with an application to California’s RIPA data on police stops, where it clarifies ambiguous results from the standard outcome test. The talk concludes by noting the test’s simplicity and its use in a federal court case involving the Philadelphia Police Department.

188 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides a valuable contribution by addressing a well-known problem in discrimination testing. The argumentation is rigorous: the speaker clearly defines the statistical model, explains the assumptions, and provides both theoretical results and empirical validation. The use of real-world data and simulations strengthens the credibility of the proposed method. The presentation is well-structured, building from simple examples to the formal theorem and its applications.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is high: the method is published in PNAS, a reputable peer-reviewed journal. The speaker cites relevant literature, including Gary Becker’s work, and discusses limitations. The title accurately reflects the content. The talk is based on original research and is presented by the author, ensuring accuracy.

128 words

Title / Content Match

The title accurately reflects the content: a presentation of a simple, statistically robust test for discrimination.

Quality & Reliability

8/10

Presentation of a peer-reviewed method (PNAS publication) with rigorous mathematical framework, empirical validation on multiple datasets, and application to real-world data. The speaker is a PhD candidate in Statistics at Harvard, and the work is co-authored with a renowned professor. The method is clearly explained, assumptions are discussed, and limitations are acknowledged.

Key Moments

Cited Sources

  • PNAS paper on robust outcome test — The speaker mentions that the work was published in PNAS last year.

Concurring Sources

Contribution & Novelties

The talk presents a novel statistical test that combines the benchmark and outcome tests to provide a more reliable indicator of discrimination. The key innovation is the formalization of the conditions under which this combined test is valid, using the monotone likelihood ratio property. The method is simple to implement and has strong statistical guarantees, making it practical for policymakers and researchers.

Pour aller plus loin :

  • Monotone likelihood ratio — The property is central to the test’s validity.
  • Becker’s outcome test — The original test proposed by Gary Becker.
  • Inframarginality — The concept explaining why outcome tests can be misleading.

101 words

Radar Profile

The radar profile shows high scores in information quality, technical level, and reliability, with a slightly lower score in information quantity due to the concise presentation. This indicates a technically rigorous and reliable talk, though not exhaustive in scope.

Reliability 8/10

💬 No comments were provided for analysis.