Shancong Mou - Derivative-Informed Training of Neural Operators on the Fly - IPAM at UCLA

Shancong Mou - Derivative-Informed Training of Neural Operators on the Fly - IPAM at UCLA

🎙 Shancong Mou 👥 42K 📅 May 22, 2026 ⏱ 39 min 👁 338 📄 expert opinion 🧭 2026-08-13
Available in: English (current) Français

Keywords

neural operatorsderivative informationJacobiansensitivity equationpreconditioning

Summary

Shancong Mou presents a method for training neural operators with derivative information generated on the fly, avoiding expensive offline data generation. The talk begins by motivating the importance of Jacobian accuracy in downstream PDE-constrained optimization tasks. It reviews existing derivative-informed training approaches that use offline-generated derivative pairs, then introduces a novel on-the-fly approach using a sketched tangent consistency loss. This loss enforces the sensitivity equation in a data-free manner, but suffers from poor conditioning. The speaker proposes preconditioning techniques, including cheap preconditioners for elliptic problems and Krylov subspace methods for challenging cases like Helmholtz. Experiments on 2D Helmholtz, 2D nonlinear diffusion, and 1D Burgers equations show that preconditioned on-the-fly training achieves accuracy comparable to offline derivative-informed training, significantly improving over standard training. The talk concludes with ongoing work and emphasizes the potential of this approach for efficient surrogate modeling.

139 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation provides valuable insights into a practical problem in neural operator training. The argumentation is well-structured, starting with motivation, then presenting the proposed method, addressing its limitations, and providing empirical evidence. The speaker clearly explains the intuition behind the sketching and preconditioning techniques, making the approach accessible. The results are promising, showing significant improvements in accuracy for the tested PDEs. However, the talk is primarily based on the speaker’s own research and lacks extensive comparison with other recent methods. The argumentation is solid but could benefit from more rigorous theoretical analysis and broader experimental validation.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, with clear problem formulation and methodology. The speaker references relevant literature on derivative-informed training and physics-informed neural networks, but does not provide specific citations during the talk. The title accurately reflects the content, focusing on derivative-informed training with on-the-fly generation. The presentation is part of a recognized workshop, adding credibility. However, the lack of explicit source citations and the reliance on the speaker’s own work limit the verifiability. The title is appropriate and does not overstate the content.

194 words

Title / Content Match

The title accurately reflects the content, focusing on derivative-informed training of neural operators with on-the-fly generation.

Quality & Reliability

8/10

Presentation by a domain expert at a recognized workshop, with clear methodology and empirical results, but limited peer-reviewed sources and no external verification.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk presents a novel method for training neural operators with derivative information without offline data generation, using a sketched tangent consistency loss and preconditioning. This approach reduces data generation costs while maintaining accuracy benefits. The method is demonstrated on several PDEs, showing significant improvements over standard training.

Pour aller plus loin :

  • Neural Operator Learning — Background on neural operators.
  • Physics-Informed Neural Networks — Related approach for incorporating physics into neural networks.
  • Preconditioning — Mathematical background on preconditioning techniques.

80 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded presentation with strong technical depth, reliable information, and good structure. The lowest score is in 'quantite_information' but still high, reflecting the focused scope of the talk.

Reliability 8/10