Borjan Geshkovski

Borjan Geshkovski

🎙 Borjan Geshkovski 👥 4K 📅 May 3, 2026 ⏱ 33 min 👁 35 📄 original study 🧭 2026-08-13
Available in: English (current) Français

Keywords

transformershomogenizationstochastic differential equationsmean-field limitover-smoothing

Summary

The talk presents a mathematical framework for analyzing transformers at initialization by deriving a homogenized limit as the layer depth becomes large. The speaker, Borjan Geshkovski, introduces a simplified transformer model without MLPs and masked attention, focusing on self-attention with random Gaussian weights. By scaling the time step appropriately, he derives a system of stochastic differential equations (SDEs) on the unit sphere, driven by common noise. The key result is a theorem showing that the discrete transformer iterates converge to the solution of this SDE in a weak sense. The talk then uses this limiting model to study over-smoothing, a phenomenon where tokens cluster together, and shows that in the centered Gaussian case, the drift vanishes, leaving only noise and a corrector term. By analyzing the SDE, the speaker derives conditions on hyperparameters (like the variance of value matrices) to avoid over-smoothing, even in regimes where the noise is scaled as 1/sqrt(L). The talk concludes with a discussion of the implications for scaling transformers and mentions ongoing work.

168 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides a novel and rigorous mathematical perspective on transformer architectures, bridging the gap between deep learning and stochastic analysis. The argumentation is solid, building from a clear model to a limiting SDE via heuristic derivations and a formal theorem. The speaker carefully explains the role of each component and the scaling assumptions, making the reasoning transparent. The value lies in offering a principled way to understand hyperparameter choices and over-smoothing, which is a known practical issue. The presentation is technical but well-structured, with a clear narrative from model to limit to application.

Scientific Rigor, Source Quality, Title Accuracy

The talk is based on a recent paper (arXiv preprint) and references prior work on stochastic modified equations and over-smoothing in transformers. The speaker does not provide explicit citations during the talk, but the description includes a link to the paper. The title is minimal, but the content is coherent. The scientific rigor is high, with a clear theorem statement and acknowledgment of heuristic steps. The sources are not explicitly listed, but the reliance on established stochastic calculus methods adds credibility. The adequacy between title and content is acceptable, though the title is not descriptive.

204 words

Title / Content Match

The title is minimal (just the speaker's name), but the content is a coherent talk on homogenized transformers, so the title is not misleading.

Quality & Reliability

8/10

Presentation of a rigorous mathematical framework for homogenizing transformers, with explicit derivations and a theorem statement. The approach is based on stochastic calculus and mean-field limits, and the speaker acknowledges the heuristic nature of some steps. The work is recent and presented at a conference, but the lack of peer review and the simplified setting (no MLPs, no masked attention) limit the immediate applicability.

Key Moments

Cited Sources

  • Homogenized transformers (arXiv preprint) — The paper presenting the homogenization framework for transformers.

Concurring Sources

Dissenting Sources

  • Attention is All You Need — The original transformer paper does not consider homogenization or stochastic limits, but provides the architectural basis.

Contribution & Novelties

The talk introduces a novel mathematical framework for analyzing transformers at initialization by deriving a homogenized limit as the number of layers grows. This provides a principled way to understand hyperparameter scaling and over-smoothing, which is a known practical issue. The approach is original in its use of stochastic calculus and mean-field limits to derive a tractable SDE model.

Pour aller plus loin :

95 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and technical level, with a slightly lower but still strong reliability score. This indicates a technically dense and reliable presentation, though the lack of explicit citations and the heuristic nature of some steps slightly reduce the reliability.

Reliability 8/10

💬 Sur les 0 commentaires analysés, aucune tendance n'est disponible.