High dimensional statistics - session 2

High dimensional statistics - session 2

🎙 Robust and Interpretable Machine Learning Lab 👥 1K 📅 October 15, 2025 ⏱ 83 min 👁 212 📄 lecture 🧭 2026-08-16
Available in: English (current) Français

Keywords

high-dimensionalBayes optimal errorplug-in classifierconvergence in probabilitysparsity

Summary

This is the second session of a course on high-dimensional statistics. The instructor begins by reviewing the classification problem from the previous session, where the optimal classifier (Bayes classifier) is based on comparing the densities of two normal classes. The Bayes error is expressed as Phi(-sqrt(mu^T Sigma^{-1} mu)). The focus then shifts to the plug-in classifier, where the unknown mean mu is estimated from data. The goal is to analyze the asymptotic behavior of the classification error of this plug-in classifier as the sample size n increases. The instructor introduces convergence in probability and the continuous mapping theorem. Through a series of algebraic manipulations and change of variables, the error is expressed in terms of a quantity Q = mu^T Sigma^{-1} mu and a random vector Z’ ~ N(0, I). The analysis reveals that the error converges to Phi(-sqrt(Q) / sqrt(Q + alpha_0)), where alpha_0 is the limit of d/n. If alpha_0 = 0 (classical regime), the error converges to the Bayes error. If alpha_0 > 0 (high-dimensional regime), the error is larger, illustrating the curse of dimensionality. The instructor discusses the need for additional assumptions, such as sparsity, to overcome this challenge. The lecture ends with a discussion of how sparsity can help in high-dimensional settings.

207 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides a rigorous derivation of the asymptotic error of a plug-in classifier in high-dimensional settings. The argumentation is solid, building from first principles of probability and statistics. The instructor carefully introduces necessary assumptions, such as the limit of d/n, and explains their implications. The value lies in clearly demonstrating the curse of dimensionality and motivating the need for structural assumptions like sparsity. The presentation is mathematically detailed, making it valuable for advanced students or researchers.

Scientific Rigor, Source Quality, Title Accuracy

The lecture is mathematically rigorous, with derivations that follow standard statistical theory. The instructor does not cite external sources, but the content is based on well-established concepts in high-dimensional statistics. The title accurately reflects the content, which is a continuation of a course on high-dimensional statistics. The presentation is clear, though some notational errors are made and corrected during the lecture, which is acceptable in a live setting.

160 words

Title / Content Match

The title accurately reflects the content, which is the second session of a course on high-dimensional statistics.

Quality & Reliability

8/10

The lecture is mathematically rigorous, deriving results step-by-step from probability theory and asymptotic statistics. The instructor clearly explains assumptions and limitations, and the content aligns with established statistical theory. However, the video is a lecture, not peer-reviewed, and the presentation has some minor notational errors that are corrected on the fly.

Key Moments

Contribution & Novelties

The lecture provides a clear and rigorous derivation of the asymptotic error of a plug-in classifier in high-dimensional settings, highlighting the role of the ratio d/n. It effectively demonstrates the curse of dimensionality and motivates the need for structural assumptions like sparsity. The presentation is pedagogical, making complex concepts accessible to advanced students.

Pour aller plus loin :

  • High-dimensional statistics — Overview of the field and its challenges.
  • Concentration inequality — Tools used to bound deviations of random variables, relevant to convergence results.
  • Sparse model — Concept of sparsity and its role in high-dimensional inference.

95 words

Radar Profile

The radar profile shows high scores in quantitative information, technical level, and reliability, reflecting the mathematically rigorous and well-structured lecture. The qualitative information score is also high, indicating the depth of explanation. The overall profile suggests a highly technical and reliable educational resource.

Reliability 8/10