Shantenu Jha - All (Foundation) Models are wrong. Some (Foundation) Models can be made useful.

Shantenu Jha - All (Foundation) Models are wrong. Some (Foundation) Models can be made useful.

🎙 Shantenu Jha 👥 42K 📅 April 17, 2026 ⏱ 42 min 👁 707 📄 expert opinion 🧭 2026-08-13
Available in: English (current) Français

Keywords

foundation modelsfusionGenesis missionHPCAI for science

Summary

Shantenu Jha, a high-performance computing engineer from Princeton Plasma Physics Lab, presents an overview of the DOE’s Genesis mission, emphasizing the concept of ‘AI advantage’ in scientific discovery. He argues that foundation models are not universally correct but are trade-offs between computational cost and accuracy, and that the optimal model depends on the computational budget. He introduces a hierarchy of models (reasoning models, foundation models, surrogates) and stresses the importance of coupling between them and with traditional simulations. He discusses the trade-off between pre-training and test-time compute, citing papers on reasoning models like DeepSeek and process reward models. He then presents the design of ‘Ignos’, a foundation model for fusion, which consists of two specialized models: Helios for experimental data and Metis for simulation data. He also discusses the Genesis platform, which aims to integrate HPC, experimental facilities, and AI models to support the creation of many foundation models across disciplines. The talk concludes with a discussion of coupling AI with HPC in various modes (‘AI in, out, and about HPC’) and the importance of the American Science Cloud.

179 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the practical considerations of building and deploying foundation models for scientific applications, particularly in fusion. The argumentation is solid, drawing on recent papers and the speaker’s direct experience with the Genesis mission. The speaker effectively makes the case that foundation models are not one-size-fits-all and that computational budget allocation is a critical design decision. The discussion of the trade-off between training and test-time compute is well-supported by references to specific research. The presentation of the Ignos architecture offers a concrete example of how these principles are applied.

Scientific Rigor, Source Quality, Title Accuracy

The speaker cites several relevant papers, including the Epokaia report commissioned by Google DeepMind, the DeepSeek paper, and work on process reward models. These are credible sources. The title is a clever adaptation of George Box’s famous quote and accurately reflects the talk’s theme. The talk is well-structured and the speaker is transparent about the preliminary nature of the results. The description provides a link to the workshop page, which is a reliable source for context.

184 words

Title / Content Match

The title accurately reflects the content, which discusses the limitations of foundation models and the need for computational budget trade-offs, echoing George Box's aphorism.

Quality & Reliability

8/10

The talk is given by a leading expert in high-performance computing at Princeton Plasma Physics Lab, presenting the DOE's Genesis mission. It references specific papers and projects, and the content is grounded in current research. However, it is an opinion/expert talk rather than a peer-reviewed study, and some claims are forward-looking.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk provides a unique perspective from a high-performance computing engineer on the challenges and opportunities of building foundation models for scientific applications, specifically fusion. It emphasizes the importance of computational budget trade-offs and the need for a hierarchy of models. The presentation of the Ignos architecture is a novel contribution, as it is an early-stage project. The talk also highlights the Genesis mission as a coordinated effort to advance AI in science.

Pour aller plus loin :

134 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and informative talk. The speaker demonstrates strong expertise and provides a balanced view of the topic, with a slight emphasis on the computational aspects.

Reliability 8/10