Aditya Varre: Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

Aditya Varre: Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

🎙 Aditya Varre 👥 3K 📅 October 26, 2025 ⏱ 48 min 👁 136 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

in-context learningtransformersn-gramstationary pointsgradient vanishing

Summary

The talk by Aditya Varre, a PhD student at EPFL, presents a theoretical analysis of why transformers exhibit plateaus during training on in-context learning tasks. The study focuses on a simplified two-layer transformer trained on data generated by an n-gram language model (a higher-order Markov chain). The authors show that sub-n-gram solutions, which only capture partial dependencies, are near-stationary points of the loss landscape. This explains the stage-wise learning dynamics observed in practice, where the model first learns simpler patterns before transitioning to more complex ones. The analysis introduces a simplified transformer architecture with concatenated attention heads and orthogonal embeddings, allowing for a tractable characterization of the gradient. They derive conditions under which the gradient vanishes, showing that when the model’s score function depends only on a partial history, the loss gradient can be zero even if the model is not the true data-generating process. The talk includes empirical evidence supporting the theoretical findings, demonstrating that transformers trained on n-gram tasks exhibit the predicted stage-wise transitions. The work provides insights into the mechanistic interpretation of in-context learning and the role of attention in capturing syntactic structures.

186 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides a valuable contribution by offering a theoretical explanation for the plateau phenomenon in transformer training, a topic that has been empirically observed but not fully understood. The argumentation is rigorous: the authors define a clear mathematical framework, derive conditions for stationary points, and validate their theory with experiments. The use of a simplified transformer architecture is justified as a tractable model that retains essential features. The presentation is well-structured, building from motivation to formal results and empirical validation. The claims are supported by a proof sketch and references to a peer-reviewed paper, enhancing credibility. The discussion of partial solutions and their role in stage-wise learning is insightful and adds depth to the understanding of in-context learning.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is high: the talk is based on a paper published at ICML 2025, a top-tier conference, and the presenter provides a clear mathematical exposition. The sources cited include the paper itself (arXiv link) and references to prior work on n-gram estimation and transformer training dynamics. The title accurately reflects the content, focusing on the specific finding that sub-n-grams are near-stationary points. The talk does not overstate its claims and acknowledges the limitations of the simplified architecture. The use of orthogonal embeddings and concatenation is a deliberate simplification, and the presenter notes that similar results can be observed empirically with learned embeddings. Overall, the sources are appropriate and the title is well-aligned with the content.

251 words

Title / Content Match

The title accurately reflects the content: the talk focuses on learning in-context n-grams with transformers and demonstrates that sub-n-grams are near-stationary points.

Quality & Reliability

8/10

The talk presents a rigorous theoretical analysis of transformer training dynamics on a synthetic n-gram task, with a clear mathematical framework and empirical validation. The claims are supported by a formal proof sketch and references to a peer-reviewed ICML paper. The presentation is coherent and the methodology is sound, though the simplified architecture and assumptions (orthogonal embeddings, concatenation) may limit direct applicability to full-scale transformers.

Key Moments

Cited Sources

Concurring Sources

Dissenting Sources

  • Attention is All You Need — The original transformer architecture uses residual connections and layer normalization, which are simplified in this work. The concatenation of heads and skip connections is non-standard, potentially limiting direct applicability.

Contribution & Novelties

The talk offers a novel theoretical framework to explain the plateau phenomenon in transformer training by identifying sub-n-gram solutions as near-stationary points. This provides a mechanistic understanding of stage-wise learning, where the model gradually adds syntactic structures. The simplified architecture allows for a rigorous analysis of the loss landscape, which is a significant contribution to the theoretical understanding of in-context learning.

Pour aller plus loin :

124 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and technical level, indicating a dense and rigorous presentation. The fiabilite_globale is also high, reflecting the solid theoretical foundation. The profile suggests a talk that is highly informative and technically deep, suitable for an expert audience.

Reliability 8/10