Why Muon Is Good but May Not Be Optimal

Why Muon Is Good but May Not Be Optimal

🎙 Weijie Su 👥 4K 📅 February 20, 2026 ⏱ 60 min 👁 308 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

Muonoptimizationpreconditioningsingular value decompositionlanguage models

Summary

The talk, presented by Weijie Su at a CRUNCH Group seminar, examines the Muon optimization method for training large language models. Muon, introduced in December 2024, updates weights along the orthogonalized gradient (U V^T) from SVD, discarding singular values. Su provides two theoretical perspectives to understand why Muon works and why it may not be optimal. First, he introduces a unifying framework distinguishing preconditioning for curvature anisotropy (Adam) and gradient anisotropy (Muon). This leads to a new method, PolarGrad, which incorporates curvature information via the nuclear norm of the gradient into the learning rate, improving convergence. Second, he presents an isotropic curvature model assuming curvature isotropy across perturbation directions. Under a general growth condition, the optimal update makes the gradient’s spectrum more homogeneous. The orthogonalized gradient becomes optimal only when curvature exhibits a phase transition in growth. These perspectives suggest that Muon’s gradient orthogonalization is directionally correct but may not be strictly optimal, and the model can guide designing new optimization methods. The talk includes Q&A and references to arXiv papers 2505.21799 and 2511.00674.

174 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the theoretical foundations of Muon, a recent optimization method. It offers two distinct perspectives: one based on preconditioning and another on isotropic curvature models. The argumentation is rigorous, with mathematical derivations and references to empirical results. The introduction of PolarGrad as a new method is a significant contribution. The speaker clearly explains the limitations of existing approaches and justifies the need for new theoretical frameworks. The discussion of learning rate adaptation via nuclear norm is particularly insightful. Overall, the talk is highly valuable for researchers in optimization and deep learning.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, with clear mathematical reasoning and references to arXiv papers. The speaker is a recognized expert, and the content aligns with his research. The title accurately reflects the content, as the talk indeed explains why Muon is good and then argues it may not be optimal. The sources cited are relevant and credible. The talk does not include any commercial or promotional content. The Q&A session adds to the rigor by addressing potential concerns.

189 words

Title / Content Match

The title accurately reflects the content: the talk explains why Muon is effective and then argues it may not be optimal, offering two theoretical perspectives.

Quality & Reliability

8/10

Talk by a recognized researcher (Weijie Su, UPenn) presenting theoretical analysis and new methods (PolarGrad) based on arXiv papers. Claims are supported by mathematical derivations and some experiments, but not peer-reviewed in this presentation.

Key Moments

Cited Sources

  • arXiv:2505.21799 — Referenced as the basis for the first perspective on preconditioning.
  • arXiv:2511.00674 — Referenced as the basis for the second perspective on isotropic curvature model.

Concurring Sources

Contribution & Novelties

The talk provides a novel theoretical framework to understand Muon’s effectiveness and limitations. It introduces PolarGrad, a new optimization method that incorporates curvature information via nuclear norm. The isotropic curvature model offers a new perspective on optimal updates. These contributions advance the theoretical understanding of matrix-based optimizers.

Pour aller plus loin :

73 words

Radar Profile

The radar profile shows high scores in technical level and information quality, with slightly lower scores in information quantity and reliability. This reflects a specialized talk with strong theoretical content but limited breadth and some reliance on unpublished results.

Reliability 8/10