Jiaoda Li: Characterizing the Expressivity of Transformer Language Models

Jiaoda Li: Characterizing the Expressivity of Transformer Language Models

🎙 Jiaoda Li 👥 3K 📅 August 18, 2025 ⏱ 40 min 👁 58 📄 original study 🧭 2026-08-17
Available in: English (current) Français

Keywords

transformersexpressivityformal languagestemporal logicLTL

Summary

The talk presents a formal characterization of the expressive power of transformer language models under fixed precision, no positional encodings, and softmax attention. The authors introduce a fragment of first-order logic with two variables, PFO2, and show it is equivalent to LTL with only past operators (LTL_past). They prove that transformers can be translated into PFO2, and conversely, LTL_past formulas can be simulated by transformers. This establishes an exact equivalence between transformers and LTL_past, which is shown to correspond to left-deterministic polynomials, a class of regular languages recognized by partially ordered DFAs. The theoretical results are validated experimentally on a set of languages predicted to be learnable or not by transformers, with results aligning with predictions. The talk also discusses implications for language modeling and internal representations.

127 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides a significant theoretical contribution by precisely characterizing the expressivity of transformers in terms of a well-understood logical fragment. The argumentation is rigorous, building on prior work and addressing technical challenges such as simulating summation with threshold counting. The proof sketches are clear and the experimental validation strengthens the claims. The discussion of limitations, such as the fixed precision assumption, adds to the credibility.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, with a clear theoretical framework and proofs. The main source is the paper on arXiv (2505.23623), which is appropriately referenced. The title accurately reflects the content. The presentation is self-contained, providing necessary background on formal languages and logic. The experimental methodology is sound, using a length generalization setting. The talk does not overstate its findings and acknowledges assumptions.

144 words

Title / Content Match

The title accurately reflects the content: the talk characterizes the expressivity of transformer language models in terms of a fragment of temporal logic.

Quality & Reliability

8/10

The talk presents a formal theoretical result with a rigorous proof sketch, supported by experimental validation. The methodology is clearly explained, and the claims are grounded in established formal language theory. The presentation is technical and precise, with appropriate caveats about assumptions.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk provides an exact characterization of the expressive power of transformers under fixed precision and softmax attention, which was previously unknown. It introduces a new logical fragment (PFO2) and shows its equivalence to LTL_past, and then to left-deterministic polynomials. This bridges formal language theory and neural network expressivity, offering a precise boundary for what transformers can and cannot learn. The experimental validation on a range of languages confirms the theoretical predictions.

Pour aller plus loin :

101 words

Radar Profile

The radar profile shows high scores in information quality, technical depth, and reliability, with slightly lower scores in information quantity and global reliability due to the narrow focus and reliance on a single source. This indicates a specialized, rigorous talk suitable for an expert audience.

Reliability 8/10