Kulin Shah: Understanding Token Ordering in Masked Diffusions

Kulin Shah: Understanding Token Ordering in Masked Diffusions

🎙 Kulin Shah 👥 3K 📅 November 17, 2025 ⏱ 48 min 👁 74 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

masked diffusionautoregressivetoken orderingscaling lawshardness

Summary

The talk by Kulin Shah discusses the role of token generation order in masked diffusion models (MDMs) compared to autoregressive (AR) language models. It begins by motivating diffusion models as efficient alternatives to AR models, citing examples like Gemini and Mercury Coder. The speaker then explains the training and inference of MDMs, highlighting that MDMs are order-agnostic, training on all permutations, whereas AR models have a left-to-right inductive bias. Through experiments on language modeling, they show that AR models achieve better scaling laws when aligned with natural left-to-right order, and that MDMs perform similarly to random permutation training. For tasks with nonlinear structure, such as coding and Sudoku, the natural order is not left-to-right, and AR training can be suboptimal. The talk presents a theoretical result showing that for a specific distribution with a natural left-to-right order, MDMs face computational hardness due to the difficulty of infilling tasks, related to planted graph coloring. Empirical evidence confirms that some infilling tasks are harder to learn. The speaker proposes an inference strategy for MDMs that selects tokens to unmask based on uncertainty, which can improve performance on tasks without a fixed order. The talk concludes that while MDMs may be less effective for natural language due to the left-to-right bias, they can be advantageous for tasks with flexible generation order, and the proposed inference method can help realize this potential.

228 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the strengths and limitations of masked diffusion models versus autoregressive models, particularly regarding token ordering. The argumentation is solid, combining theoretical results with empirical evidence. The theoretical hardness result is well-motivated and clearly explained, linking MDM training to planted graph coloring. The empirical scaling law experiments are convincing and support the claim that MDMs underperform AR models on natural language due to the lack of left-to-right inductive bias. The proposed uncertainty-based inference is a practical contribution that addresses the identified limitation. The speaker also discusses the importance of generation order in tasks like coding and Sudoku, providing a broader perspective. Overall, the talk is well-structured and presents a balanced view, acknowledging both the potential and the challenges of diffusion models.

Scientific Rigor, Source Quality, Title Accuracy

The talk is based on two papers: one published at ICML 2024 (with Jan, Vasilis, Sham, and Sitan) and another from 2023 (with Nishan, Jin, and Rein). The speaker references the paper link in the description (https://arxiv.org/abs/2502.06768) . The presentation is rigorous, with clear definitions and proofs sketched. The title accurately reflects the content. The speaker does not mention any external sources beyond the papers, but the theoretical and empirical evidence is presented in a self-contained manner. The talk is suitable for a technical audience familiar with diffusion models and language modeling.

233 words

Title / Content Match

The title accurately reflects the content, focusing on token ordering in masked diffusion models.

Quality & Reliability

8/10

The talk is based on two peer-reviewed papers (ICML 2024 and 2023) and presents theoretical results with proofs and empirical evidence. The speaker is a PhD student at UT Austin, and the presentation is clear and rigorous. However, the talk is a seminar presentation, not a peer-reviewed publication itself, and some claims rely on unpublished or pre-print results.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk provides a novel analysis of token ordering in masked diffusion models, highlighting the importance of generation order for performance. It introduces a theoretical hardness result linking MDM training to planted graph coloring, and proposes an uncertainty-based inference strategy to mitigate the issue. The talk also presents empirical scaling law comparisons between AR and MDM models, showing that MDMs underperform on natural language due to the lack of left-to-right bias.

Pour aller plus loin :

102 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and technical level, with a slightly lower but still strong reliability score. This indicates a technically dense and informative talk with solid theoretical and empirical backing.

Reliability 8/10