Optimal Implicit and Explicit Regularization in High Dimensional Continual Linear Regression

Optimal Implicit and Explicit Regularization in High Dimensional Continual Linear Regression

🎙 Gilad Karpel 👥 385 📅 April 16, 2026 ⏱ 59 min 👁 71 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

continual learningL2 regularizationimplicit regularizationgradient descentgeneralization

Summary

Gilad Karpel presents his research on optimal regularization in high-dimensional continual linear regression. He addresses the problem of catastrophic forgetting in continual learning, where a model learns sequentially from tasks. He focuses on regularization-based methods, specifically L2 regularization towards the previous predictor. He derives the optimal regularization strength scaling as T/ln T with the number of tasks T under i.i.d. teachers. He also shows a close relationship between early-stopped gradient descent and L2 regularization, providing a cheaper alternative with similar performance. He identifies the optimal step size scaling as ln T/(NT). He extends the analysis to diagonal regularization schemes and provides conditions for L2 optimality. He also presents a counterexample where the optimal regularization remains finite for non-i.i.d. teachers. The theoretical findings are validated through experiments on linear regression and neural networks.

132 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides significant theoretical insights into continual learning, offering exact scaling laws for regularization strength. The argumentation is rigorous, with clear assumptions and mathematical derivations. The speaker motivates the problem well and connects theory to practice. The empirical validation on synthetic and real data strengthens the claims. The presentation is well-structured and accessible to a technical audience.

Scientific Rigor, Source Quality, Title Accuracy

The speaker cites prior work in continual learning and high-dimensional statistics, though specific references are not detailed in the talk. The title accurately reflects the content. The talk is based on two papers, one accepted at a top conference and another submitted. The methodology appears sound, with clear statistical assumptions and proofs. The empirical results support the theoretical findings.

132 words

Title / Content Match

The title accurately reflects the content, focusing on optimal regularization in high-dimensional continual linear regression.

Quality & Reliability

8/10

The talk presents rigorous theoretical results with proofs, validated by experiments on synthetic and real data. The speaker is a master's student at Technion, and the work is accepted at a top conference. The presentation is clear and includes mathematical derivations.

Key Moments

Cited Sources

  • Optimal Regularization in High Dimensional Continual Linear Regression — The paper presented in the talk, accepted at a previous AI conference.
  • Optimal Implicit and Explicit Regularization in High Dimensional Continual Linear Regression — The second paper, recently submitted, also presented in the talk.

Concurring Sources

Contribution & Novelties

The talk provides novel theoretical results on optimal regularization in continual learning, specifically deriving scaling laws for L2 regularization and early stopping. It offers practical guidance for setting regularization strength. The work bridges theory and practice, with experiments validating the theory.

Pour aller plus loin :

93 words

Radar Profile

The radar profile shows high scores in technical level and information quality, with slightly lower scores in quantity and reliability, reflecting the specialized nature of the talk and the reliance on theoretical results.

Reliability 8/10