How abundant are good interpolators?

How abundant are good interpolators?

🎙 Ahmed El Alaoui 👥 75K 📅 August 7, 2026 ⏱ 44 min 👁 326 📄 original study 🧭 2026-08-08
Available in: English (current) Français

Keywords

interpolatorsgeneralization errorlarge deviation principlelinear classificationoverparameterization

Summary

The talk addresses the generalization performance of interpolating classifiers in high-dimensional settings. The speaker considers linear classifiers on labeled data, where an interpolator is a classifier that correctly classifies all training points with a given margin (possibly negative). Under two data-generating models (logistic and Gaussian mixture), he establishes a large deviation principle for the proportion of interpolators achieving a given overlap with the true signal direction. This yields a precise description of the distribution of generalization errors among interpolators. The typical interpolator is shown to have a generalization error inferior to that of ERM by gradient descent, indicating that most interpolators are ‘bad’. The talk also discusses the interpolation threshold (capacity) and the geometry of the set of interpolators, contrasting convex (positive margin) and non-convex (negative margin) regimes. The results are asymptotic, with n and d growing proportionally. The speaker concludes by mentioning open problems, such as sampling typical interpolators and understanding algorithmic access.

154 words

Critical Evaluation

The talk presents a rigorous theoretical analysis of the generalization properties of interpolators in high-dimensional linear classification. The speaker clearly defines the problem, introduces the relevant models, and states a main theorem (large deviation principle) that provides a complete characterization of the proportion of interpolators with a given generalization error. The mathematical framework is sound, relying on established tools from high-dimensional probability and statistical physics. The assumptions are clearly stated, and the asymptotic regime (n,d → ∞ with n/d → α) is standard in the field. The results are novel and contribute to the understanding of the interpolation regime, a topic of significant recent interest. The speaker also discusses the interpolation threshold, connecting to prior work, and highlights the contrast between convex and non-convex regimes. The presentation is well-structured, with clear motivation and intuitive explanations of the geometric picture. However, the talk is highly technical and may be inaccessible to non-specialists. The speaker does not provide full proofs, but sketches the main ideas. The results are limited to specific data distributions (logistic and Gaussian mixture), and the extension to more general settings is not discussed. The claim that typical interpolators are ‘bad’ is based on the comparison with ERM, but the practical implications are not explored. Overall, the talk is of high scientific quality, with clear contributions to the theory of interpolation in machine learning. The title accurately reflects the content, and the presentation is coherent. The only minor weakness is the lack of discussion of practical implications or connections to other work beyond the cited references.

257 words

Title / Content Match

The title accurately reflects the central question: the abundance (proportion) of interpolators with good generalization performance.

Quality & Reliability

8/10

The talk presents rigorous mathematical results (large deviation principles) with clear assumptions and proofs sketched. The speaker is a recognized researcher at Cornell, and the work is joint with a student. The presentation is technical and precise, with no obvious overclaims. However, the results are asymptotic and rely on idealized models, and the talk does not provide full proofs or external validation within the video.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk provides a novel large deviation principle for the proportion of interpolators achieving a given generalization error in high-dimensional linear classification. This goes beyond prior work that focused on the existence and capacity of interpolators, offering a precise characterization of the distribution of generalization errors within the set of interpolators. The finding that typical interpolators are ‘bad’ (inferior to ERM) is a significant contribution to the debate on benign overfitting.

Pour aller plus loin :

  • Large deviations theory — Provides the mathematical foundation for the main results.
  • Benign overfitting — Discusses the phenomenon where interpolating models generalize well, contrasting with the talk’s findings.
  • High-dimensional statistics — Context for the asymptotic regime and challenges.
  • Statistical physics of learning — Related to the replica method and typical behavior in high-dimensional models.

130 words

Radar Profile

The radar profile shows high scores in quality of information and technical level, reflecting the rigorous mathematical nature of the talk. The quantity of information is also high, but the global reliability is slightly lower due to the reliance on idealized models and the lack of full proofs in the presentation.

Reliability 8/10