
Shantenu Jha - All (Foundation) Models are wrong. Some (Foundation) Models can be made useful.
Keywords
Summary
179 words
Critical Evaluation
Value of the Information & Strength of the Argument
The talk provides valuable insights into the practical considerations of building and deploying foundation models for scientific applications, particularly in fusion. The argumentation is solid, drawing on recent papers and the speaker’s direct experience with the Genesis mission. The speaker effectively makes the case that foundation models are not one-size-fits-all and that computational budget allocation is a critical design decision. The discussion of the trade-off between training and test-time compute is well-supported by references to specific research. The presentation of the Ignos architecture offers a concrete example of how these principles are applied.
Scientific Rigor, Source Quality, Title Accuracy
The speaker cites several relevant papers, including the Epokaia report commissioned by Google DeepMind, the DeepSeek paper, and work on process reward models. These are credible sources. The title is a clever adaptation of George Box’s famous quote and accurately reflects the talk’s theme. The talk is well-structured and the speaker is transparent about the preliminary nature of the results. The description provides a link to the workshop page, which is a reliable source for context.
184 words
Title / Content Match
The title accurately reflects the content, which discusses the limitations of foundation models and the need for computational budget trade-offs, echoing George Box's aphorism.
Quality & Reliability
8/10
The talk is given by a leading expert in high-performance computing at Princeton Plasma Physics Lab, presenting the DOE's Genesis mission. It references specific papers and projects, and the content is grounded in current research. However, it is an opinion/expert talk rather than a peer-reviewed study, and some claims are forward-looking.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction to the talk and the Genesis mission, emphasizing AI advantage.
- Overview of the Genesis mission: 26-27 challenge problems and the platform.
- Discussion of the Epokaia report on AI projections to 2030, highlighting the growth of compute and the shift to inference costs.
- Definition of foundation models and distinction from surrogates and reasoning models.
- Discussion of the trade-off between pre-training and test-time compute, citing DeepSeek and process reward models.
- Introduction of the Ignos foundation model for fusion, with two specialized models: Helios and Metis.
- Discussion of coupling AI with HPC: AI in, out, and about HPC.
- Description of the Genesis platform and the American Science Cloud.
- Conclusion and call for collaboration on the Genesis mission.
Cited Sources
- IPAM Workshop: Learning Models from Data for Multi-Fidelity Fusion Plasma Physics — Workshop page where the talk was recorded, providing context and related materials.
Concurring Sources
- IPAM Workshop: Learning Models from Data for Multi-Fidelity Fusion Plasma Physics — The workshop page confirms the talk's context and provides additional resources.
Contribution & Novelties
The talk provides a unique perspective from a high-performance computing engineer on the challenges and opportunities of building foundation models for scientific applications, specifically fusion. It emphasizes the importance of computational budget trade-offs and the need for a hierarchy of models. The presentation of the Ignos architecture is a novel contribution, as it is an early-stage project. The talk also highlights the Genesis mission as a coordinated effort to advance AI in science.
Pour aller plus loin :
- Foundation Models for Scientific Discovery — A survey of foundation models in science.
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — The paper cited in the talk on reasoning models.
- Process Reward Models — A paper on process reward models for reasoning.
- AlphaFold 3 — The paper on AlphaFold 3, which uses specialized models.
134 words
Radar Profile
The radar profile shows high scores across all dimensions, indicating a well-rounded and informative talk. The speaker demonstrates strong expertise and provides a balanced view of the topic, with a slight emphasis on the computational aspects.