Strange Geometric Shapes Found Inside AIs — Tom McGrath

Strange Geometric Shapes Found Inside AIs — Tom McGrath

🎙 Tom McGrath 👥 219K 📅 September 2, 2026 ⏱ 100 min 👁 23K 📄 expert opinion 🧭 2026-09-04
Available in: English (current) Français

Keywords

interpretabilityneural geometrysparse autoencodersactivation steeringreward hacking

Summary

In this episode of Machine Learning Street Talk, Tom McGrath, co-founder of Goodfire and former DeepMind researcher, discusses the state and future of AI interpretability. He argues that interpretability should be viewed as a natural science, accelerated by AI agents, and that it can enable ‘intentional design’—using interpretability as a control loop in the training process. The conversation covers how AlphaZero’s learned representations converge on human-like concepts, the challenges of steering models without falling off the data manifold, and the potential of sparse autoencoders (SAEs) to reveal interpretable features. McGrath also addresses the risks of using interpretability for training, the ‘forbidden method’ of using SAE features as rewards, and the phenomenon of models recognizing hallucinations or reward hacks yet still producing them. The episode concludes with a discussion on the limitations of SAEs, which may fracture higher-dimensional structures, and the need for new methods to capture the true geometry of neural networks.

152 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides substantial value by offering an insider perspective on cutting-edge interpretability research, with concrete examples from recent papers and practical demonstrations (e.g., the pirate-style math generalization). McGrath’s arguments are well-structured, moving from the motivation (accelerating science) to specific techniques (SAE-based gradient attribution) and their implications (intentional design, safety concerns). He acknowledges counterarguments, such as the ‘forbidden method’ concerns, and engages with them thoughtfully, strengthening the overall argumentation.

Scientific Rigor, Source Quality, Title Accuracy

The discussion is rigorous, referencing multiple arXiv papers and Goodfire research, all of which are listed in the description. The speaker’s expertise is evident, and he is careful to distinguish between established findings and speculative ideas. The title accurately reflects the content, focusing on the geometric structures found in neural networks. The video includes a brief sponsorship mention, but it does not detract from the scientific content.

151 words

Title / Content Match

The title accurately reflects the content, which focuses on the geometric structures (manifolds, features) discovered inside neural networks through interpretability research.

Quality & Reliability

8/10

The discussion is led by a researcher with deep expertise in interpretability, referencing multiple peer-reviewed papers and providing concrete examples. While the format is conversational and exploratory, the claims are grounded in published research and the speaker's direct experience, lending high credibility.

Chapters

Cited Sources

Concurring Sources

Dissenting Sources

External References

Contribution & Novelties

The video offers a compelling vision for using interpretability not just as a post-hoc analysis tool but as an integral part of the training loop, enabling ‘intentional design’ of models. It highlights the potential for AI-driven scientific discovery through interpretability and discusses the limitations of current SAE-based methods, suggesting a need for new approaches to capture the true geometry of neural networks.

Pour aller plus loin :

89 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a technically deep, reliable, and information-rich discussion. The slightly lower score in 'quantite_information' reflects the conversational format, but the density of ideas remains high.

Reliability 8/10