How to Build a Modern Diffusion Language Model

How to Build a Modern Diffusion Language Model

🎙 Volodymyr Kuleshov 👥 75K 📅 August 6, 2026 ⏱ 48 min 👁 677 📄 expert opinion 🧭 2026-08-06
Available in: English (current) Français

Keywords

diffusion language modelsmasked diffusioniterative refinementcontrollable generationinference-time scaling

Summary

Volodymyr Kuleshov from Cornell University presents an overview of modern diffusion language models, focusing on the research advances that have enabled their recent success. He begins by contrasting autoregressive models, which are inherently sequential and bottlenecked during post-training and inference-time scaling, with diffusion models that generate tokens in parallel. The talk introduces the core concept of masked diffusion, where noise is introduced by masking tokens, and the model is trained to unmask them. He explains the forward and reverse processes, drawing parallels to Gaussian diffusion used in image generation. Kuleshov then discusses extensions of simple masked diffusion, including techniques for iterative refinement, controllable generation, and efficient inference with variable-length sequences. He highlights the advantages of diffusion LLMs, such as fast generation (over 1000 tokens per second), new forms of inference-time scaling through repeated refinement, and fine-grained control, which are particularly valuable in scientific applications. The talk references leading open-source models like Nemotron Diffusion, Gemma Diffusion, and Mercury 2, and positions diffusion as a promising direction for scaling language models beyond autoregressive limitations.

172 words

Critical Evaluation

The talk provides a valuable and accessible overview of diffusion language models, a rapidly evolving area. Kuleshov, a key researcher in the field, offers credible insights into the motivations and techniques behind these models. The argumentation is clear: he identifies the sequential bottleneck of autoregressive models and positions diffusion as a parallel alternative that can unlock faster generation and new scaling axes. The technical content is solid, covering masked diffusion as a foundational approach and touching on advanced extensions like iterative refinement and controllable generation. However, the talk is a high-level survey rather than a deep dive; many claims are presented without detailed derivations or empirical evidence. For instance, the speed benchmark of >1000 tokens per second is mentioned but not contextualized with specific model configurations or hardware. The speaker also references several models (Mercury, Gemma Diffusion, Nemotron Diffusion) but does not provide comparative performance metrics or rigorous evaluation. The sources cited are limited to the Simons Institute talk page, which lacks detailed references. The talk’s strength lies in its clear conceptual framework and its relevance to current research, but it would benefit from more concrete examples and citations to primary literature. The title accurately reflects the content, and the talk is well-structured, making it a useful introduction for researchers and practitioners. The presence of a Q&A segment adds interactivity, though the audience question about token flipping is only partially addressed. Overall, the talk is informative and credible, but its reliance on anecdotal evidence and lack of detailed citations prevent it from achieving the highest rigor.

256 words

Title / Content Match

The title accurately reflects the content: the talk provides a step-by-step overview of building modern diffusion language models, from basic masked diffusion to advanced techniques used in state-of-the-art models.

Quality & Reliability

8/10

The talk is given by a leading researcher (Volodymyr Kuleshov) at a prestigious venue (Simons Institute), and presents technical content grounded in recent research and open-source models. The speaker is directly involved in the development of Mercury, a commercial diffusion LLM, lending credibility. However, the talk is a high-level overview without detailed derivations or peer-reviewed citations, and some claims (e.g., speed benchmarks) are not fully substantiated.

Key Moments

Cited Sources

Concurring Sources

Dissenting Sources

  • Autoregressive models remain competitive — Some research suggests that autoregressive models, when scaled appropriately, can still match or exceed diffusion models in certain benchmarks, challenging the talk's implicit assumption that diffusion is superior. No specific source provided.

Contribution & Novelties

The talk provides a clear, up-to-date overview of diffusion language models, synthesizing recent research advances and connecting them to practical open-source implementations. It highlights the shift from autoregressive to diffusion-based generation as a means to overcome sequential bottlenecks, and emphasizes new capabilities such as fast generation, iterative refinement, and controllable generation. The speaker’s direct involvement in developing Mercury adds practical insight.

Pour aller plus loin :

131 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and reliable presentation. The talk excels in technical depth and information quality, with slightly lower scores in novelty and source diversity, reflecting its nature as an overview rather than a groundbreaking contribution.

Reliability 8/10