Yash Sarrof: Length Generalization of Transformers with a Growing Test time Alphabet

Yash Sarrof: Length Generalization of Transformers with a Growing Test time Alphabet

🎙 Yash Sarrof 👥 3K 📅 July 23, 2026 ⏱ 44 min 👁 33 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

length generalizationtransformersC-RASPplanningvocabulary growth

Summary

The talk by Yash Sarrof, a PhD student at Saarland University, addresses the problem of length generalization in transformers, specifically when the vocabulary (alphabet) also grows at test time. He introduces a new RASP variant called C-star RASP, which extends C-RASP to handle a growing alphabet. The motivation comes from planning tasks, where generalization requires handling both longer plans and more objects. The main result is a theoretical guarantee that if a task can be expressed in C-star RASP, then transformers will generalize to both longer sequences and larger vocabularies. The proof sketch follows the framework of Huang et al., using a symbolic limit transformer that allows infinite positions and infinite symbols, with constraints of translation invariance and locality. The talk also discusses the limitations and potential applications of this work.

131 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the theoretical foundations of transformer generalization, extending prior work to a more realistic scenario where vocabulary size is not fixed. The argumentation is solid, building on established concepts like C-RASP and limit transformers, and clearly motivates the need for symbolic generalization. The speaker effectively uses examples from planning domains to illustrate the concepts. The proof sketch is coherent, though high-level, and the speaker acknowledges limitations, such as the removal of positional predicates for technical reasons.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor by referencing relevant literature, including the original RASP paper, C-RASP, and the work of Huang et al. The speaker clearly distinguishes between empirical observations and theoretical results. The title accurately reflects the content, focusing on length generalization with a growing alphabet. The sources cited are appropriate and directly related to the topic. The talk is a presentation of original research, and the speaker is transparent about the scope and limitations.

171 words

Title / Content Match

The title accurately reflects the content: the talk focuses on length generalization of transformers, specifically addressing the novel aspect of a growing test-time alphabet.

Quality & Reliability

8/10

The talk presents original research with a formal proof sketch, building on established theoretical frameworks (C-RASP, limit transformers). The speaker is a PhD student at a recognized institution, and the work is accepted at ICML. The presentation is rigorous, with clear definitions and a structured argument. However, the talk is a seminar presentation, not a peer-reviewed paper, and the proof sketch is high-level, omitting technical details.

Key Moments

Cited Sources

  • Paper 1 — Referenced as the paper on C-star RASP and length generalization with growing alphabet.
  • Paper 2 — Referenced as the paper on understanding transformers' ability to do plan verification.

Concurring Sources

  • Huang et al. ICLR paper on length generalization — Referenced in the talk as the formalization of the link between C-RASP and length generalization.

Contribution & Novelties

The talk presents a novel extension of C-RASP to handle growing alphabets, addressing a gap in existing theoretical frameworks. This is significant for planning tasks and for understanding transformer generalization in realistic settings where vocabulary size is large. The introduction of C-star RASP and the symbolic limit transformer provides a new tool for analyzing and designing transformers.

Pour aller plus loin :

98 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a technically rigorous and well-sourced presentation. The talk is particularly strong in technical depth and information quality, with slightly lower scores in information quantity due to the focused scope.

Reliability 8/10