Some Very Old and Very New Problems in Learning Theory

Some Very Old and Very New Problems in Learning Theory

🎙 Daniel Barzilai 👥 385 📅 June 12, 2026 ⏱ 38 min 👁 108 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

model collapsesynthetic dataSGDlower boundsmaximum likelihood estimation

Summary

Daniel Barzilai presents two theoretical works in learning theory. The first part addresses model collapse, the phenomenon where models trained on data contaminated by previous models degrade. Under classical regularity conditions (smoothness, constant support), the author shows that with accumulating data, model collapse does not occur even after exponentially many iterations, with error bounded by a constant times the single-iteration MLE error. However, without such assumptions, collapse can occur arbitrarily quickly. The proof relies on a Taylor expansion showing the gap between successive models scales as 1/(nT^2), summing to a constant. The second part investigates why SGD can fail on problems that are easy for other methods, such as parity. The author proposes a new approach to lower bounds for SGD, focusing on the difficulty of aligning with the important subspace in multi-index models. He shows that under certain conditions on the gradient noise, SGD requires a long time to find the relevant directions, recovering known hardness results and clarifying when heuristic SQ bounds may be misleading.

167 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into two important theoretical questions. For model collapse, the positive result under classical assumptions is reassuring, while the negative results highlight the necessity of these assumptions. The argumentation is rigorous, with clear proof sketches and appropriate caveats about the gap between positive and negative results. For SGD lower bounds, the new approach offers a more direct analysis than SQ bounds, potentially leading to more accurate hardness results. The speaker effectively motivates the problems and explains the key ideas without oversimplifying.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, with clear definitions and assumptions. The speaker cites relevant literature (e.g., Shumailov et al. on model collapse) and discusses prior work. The title accurately reflects the content, covering both a new problem (model collapse) and an old one (SGD limitations). The presentation is well-structured, and the speaker acknowledges limitations and open questions. No comments were provided for analysis.

163 words

Title / Content Match

The title accurately reflects the content: the talk addresses both a very new problem (model collapse) and a very old one (limitations of gradient-based learning).

Quality & Reliability

8/10

The talk presents rigorous theoretical results from two papers, with clear assumptions and proofs sketched. The speaker is a PhD student at Weizmann, co-advised by prominent researchers. The content is technical and precise, with appropriate caveats about limitations and gaps.

Key Moments

Cited Sources

  • Shumailov et al. on model collapse — Mentioned as early experimental work on model collapse.

Concurring Sources

  • Shumailov et al. on model collapse — Early experimental evidence of model collapse.

Contribution & Novelties

The talk presents novel theoretical results on model collapse and SGD lower bounds. For model collapse, it provides a rigorous analysis under classical assumptions, showing that collapse does not occur with accumulating data, and constructs adversarial examples where it does. For SGD, it introduces a new technique for proving lower bounds that directly analyzes the dynamics of SGD, potentially overcoming limitations of SQ bounds. This contributes to a deeper understanding of when synthetic data feedback loops are dangerous and when SGD fails.

Pour aller plus loin :

109 words

Radar Profile

The radar profile shows high scores in quality of information, technical level, and reliability, with slightly lower scores in quantity of information due to the focused scope. This indicates a technically rigorous talk with strong theoretical contributions.

Reliability 8/10