The Expressive Power of Large Language Models

The Expressive Power of Large Language Models

🎙 Gabriel Peyré 👥 79K 📅 June 12, 2026 ⏱ 61 min 👁 4K 📄 research talk 🧭 2026-08-02
Available in: English (current) Français

Keywords

LLMapproximationattentiontransformersinfinite-dimensional

Summary

Gabriel Peyré’s talk at IHES explores the mathematical foundations of large language models (LLMs), focusing on their expressive power through the lens of approximation theory. He emphasizes the challenge of handling very large contexts, which requires studying functions on infinite-dimensional spaces of token distributions. The talk breaks down the transformer architecture into three key components: normalization, attention, and MLP layers, highlighting that attention is the only layer where tokens interact. Peyré compares approximation results for perceptrons and attention mechanisms, noting that attention’s approximation properties are less understood. He discusses the role of softmax in attention, which allows focusing on specific tokens, and raises open questions about the behavior as context and layers grow. The talk is aimed at mathematicians, urging them to engage with AI research. Peyré also mentions the practical implications for coding and mathematics, where large context handling is a bottleneck. He concludes by suggesting that idealized limits of infinite context lead to probability distributions over token spaces, opening new mathematical avenues.

164 words

Critical Evaluation

The talk provides a rigorous mathematical perspective on LLMs, a topic often dominated by empirical approaches. Peyré successfully bridges approximation theory and modern AI, offering a clear decomposition of transformer layers and their roles. The emphasis on attention as the sole interaction layer is insightful, and the comparison with MLP approximation results is valuable. However, the talk is more of a research agenda than a comprehensive review; many claims are motivational and lack detailed proofs. The discussion of infinite-dimensional limits is intriguing but remains abstract, and the practical implications for training are deliberately omitted. The sources cited are minimal, with only the Carmin.tv platform mentioned, which limits the verifiability of specific claims. The title is accurate, and the content is technically deep, suitable for a mathematically inclined audience. The talk’s strength lies in framing open questions and providing a conceptual toolkit, but it could benefit from more concrete examples or references to existing literature. Overall, it is a thought-provoking contribution that stimulates further research, though it does not provide definitive answers.

171 words

Title / Content Match

The title accurately reflects the content, focusing on the expressive power of LLMs through approximation theory.

Quality & Reliability

8/10

Talk by a recognized researcher (CNRS, ENS) presenting mathematical perspectives on LLMs, with rigorous formalism and open questions. However, it is a research talk without peer-reviewed publication details, and some claims are motivational.

Key Moments

Cited Sources

  • Carmin.tv — Platform hosting the video and related scientific content.

Concurring Sources

Contribution & Novelties

This talk contributes a mathematical framework for analyzing the expressive power of LLMs, particularly focusing on attention mechanisms and their approximation properties in infinite-dimensional spaces. It highlights open questions that connect LLMs with approximation theory, offering a fresh perspective for mathematicians. The talk also emphasizes the practical bottleneck of large contexts, motivating further theoretical research.

Pour aller plus loin :

91 words

Radar Profile

The radar profile shows high scores in technical level and information quality, indicating a mathematically rigorous talk. The lower score in information quantity reflects the focus on conceptual frameworks rather than exhaustive coverage. Overall, the talk is well-balanced for a specialized audience.

Reliability 8/10

💬 Sur les 0 commentaires analysés, aucune tendance n'est disponible.