Cai Zhou: Coevolutionary Continuous Discrete DLMs, Semantic Scale Prediction via Hierarchical DLMs

Cai Zhou: Coevolutionary Continuous Discrete DLMs, Semantic Scale Prediction via Hierarchical DLMs

🎙 Cai Zhou 👥 3K 📅 January 31, 2026 ⏱ 46 min 👁 141 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

diffusion language modelshierarchical diffusioncontinuous-discrete diffusionexpressivityCTMC

Summary

The talk by Cai Zhou, a second-year PhD student at MIT, presents new modeling paradigms for diffusion language models (DLMs). It begins by contrasting autoregressive models with diffusion models, highlighting the limitations of mask diffusion models, which only have two states (mask and clean) and thus suffer from information loss and limited expressivity. The first part introduces Hierarchical Diffusion Language Models (HDLM), which add intermediate ‘cluster’ tokens between mask and clean tokens, enabling multi-level semantic prediction. The forward process uses a block-diagonal transition matrix in a continuous-time Markov chain (CTMC), and the training loss is a weighted sum of cross-entropy losses at different hierarchy levels. Experiments show that HDLM outperforms standard MDLM on perplexity, with an optimal cluster size around the square root of the vocabulary size. The second part explores continuous diffusion models, which operate on continuous representations rather than discrete tokens. The speaker analyzes the expressivity of continuous versus discrete diffusion, showing that continuous diffusion is strictly more expressive but harder to optimize. To combine advantages, they propose Coevolutionary Continuous-Discrete Diffusion (CCDD), which jointly evolves discrete and continuous representations. The talk concludes with theoretical and empirical insights, emphasizing the potential of these new paradigms for more efficient and expressive language modeling.

203 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides substantial value by addressing fundamental limitations of current diffusion language models and proposing novel solutions. The argumentation is solid, grounded in theoretical analysis (expressivity, computational complexity) and supported by experimental results. The speaker clearly motivates the need for more expressive hidden states and demonstrates how hierarchical and continuous approaches can bridge the gap. The theoretical contributions, such as proving the expressivity advantages of continuous diffusion over discrete diffusion and loop transformers, are significant. The practical challenges of continuous diffusion are honestly acknowledged, and the proposed CCDD model is a principled attempt to combine the strengths of both paradigms. The argumentation is coherent and well-structured, with clear logical flow from problem identification to solution proposal and validation.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates high scientific rigor, with detailed mathematical formulations and references to prior work (e.g., loop transformers, coconut). The speaker cites their own papers (accepted at NeurIPS 2025) and mentions collaborations with MIT, Microsoft, and other institutions. The sources are credible, though the talk itself is a presentation of original research rather than a review. The title accurately reflects the content, covering both hierarchical and coevolutionary continuous-discrete diffusion models. The presentation is technical and assumes familiarity with diffusion models and language modeling, but it is well-organized and clear. No comments were provided, so no analysis of public reception is possible.

235 words

Title / Content Match

The title accurately reflects the content, covering both hierarchical diffusion language models and coevolutionary continuous-discrete diffusion models.

Quality & Reliability

8/10

The talk presents original research from two papers accepted at NeurIPS 2025, with rigorous theoretical analysis and experimental validation. The speaker is a PhD student at MIT, and the work involves collaborations with Microsoft and other institutions. The presentation is technical and detailed, providing mathematical formulations and empirical results. However, as a seminar talk, it lacks peer-reviewed publication details and full experimental reproducibility, and the claims are based on the speaker's own work, which may have inherent biases.

Key Moments

Cited Sources

  • Semantic Scale Prediction via Hierarchical Diffusion Language Models (NeurIPS 2025) — First paper presented, introducing HDLM.
  • Coevolutionary Continuous Discrete Diffusion Models (NeurIPS 2025) — Second paper presented, introducing CCDD.

Concurring Sources

Contribution & Novelties

The talk presents two novel modeling paradigms for diffusion language models: Hierarchical Diffusion Language Models (HDLM) and Coevolutionary Continuous-Discrete Diffusion (CCDD). HDLM introduces intermediate cluster tokens to enable multi-level semantic prediction, improving expressivity and performance over standard mask diffusion. CCDD combines discrete and continuous diffusion to leverage the strengths of both, addressing the optimization challenges of continuous diffusion while maintaining expressivity. These contributions advance the theoretical understanding of diffusion models and offer practical improvements for language generation.

Pour aller plus loin :

109 words

Radar Profile

The radar profile shows high scores in quantity and quality of information, technical level, and global reliability, indicating a dense and rigorous presentation. The relatively lower score in 'fiabilite_globale' compared to others may reflect the reliance on unpublished or recently accepted work, but overall the talk is highly informative and technically sound.

Reliability 8/10