When AI Discovers the Next Transformer — Robert Lange

When AI Discovers the Next Transformer — Robert Lange

🎙 Robert Lange 👥 218K 📅 March 13, 2026 ⏱ 78 min 👁 29K 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

ShinkaEvolveAlphaEvolveco-evolutionquality-diversityAI Scientist

Summary

In this interview, Robert Lange, founding researcher at Sakana AI, discusses the ShinkaEvolve framework, which combines large language models (LLMs) with evolutionary algorithms for open-ended program search. The conversation begins with the limitations of systems like AlphaEvolve, which optimize solutions to fixed problems but require human-provided problems. ShinkaEvolve aims to co-evolve problems and solutions, drawing inspiration from POET, PowerPlay, and MAP-Elites. The architecture uses an archive of programs organized into islands, with LLMs as mutation operators and a UCB bandit to adaptively select among frontier models. Concrete results include state-of-the-art circle packing with fewer evaluations, second place in an AtCoder challenge, and evolved loss functions for mixture-of-experts models. The discussion addresses whether these systems truly think outside the box or are parasitic on starting conditions, with Lange defending the stepping-stone argument. The AI Scientist question is explored, acknowledging that current systems are more co-pilots than autonomous researchers. Lange predicts fundamental transformation of scientific research in 5-20 years, and Tim proposes the thought experiment of alien mathematical artifacts. The episode also touches on the future of work and the potential for AI to discover new architectures.

185 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video offers valuable insights into the intersection of LLMs and evolutionary computation, presenting a novel framework (ShinkaEvolve) with concrete results. The argumentation is strong, as Lange provides technical details and addresses counterarguments, such as the parasitic nature of LLMs. He defends the stepping-stone argument, citing examples from the paper. The discussion is balanced, acknowledging limitations and open questions.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is high, with references to multiple arXiv papers and a book. The sources are relevant and credible. The title accurately reflects the content, focusing on AI’s potential to discover new architectures. The discussion is well-structured, with clear explanations of technical concepts. No significant discrepancies between title and content.

125 words

Title / Content Match

The title accurately reflects the central theme: the potential for AI to autonomously discover novel architectures, exemplified by the ShinkaEvolve framework.

Quality & Reliability

8/10

The discussion is grounded in a recent preprint (ShinkaEvolve) and references several peer-reviewed or widely recognized works (POET, MAP-Elites, PowerPlay). The speaker is a founding researcher at Sakana AI, providing expert insight. However, the content is largely conversational and opinion-based, with limited critical examination of limitations.

Chapters

Cited Sources

Concurring Sources

Dissenting Sources

  • Why Greatness Cannot Be Planned — While the book argues for open-endedness without predefined objectives, ShinkaEvolve still optimizes for a given problem, which may be seen as a tension.

External References

Contribution & Novelties

The video provides an in-depth look at ShinkaEvolve, a novel framework that co-evolves problems and solutions using LLMs and evolutionary algorithms. It highlights the importance of sample efficiency and open-endedness in AI-driven discovery. The discussion offers valuable insights into the challenges and potential of autonomous research systems.

Pour aller plus loin :

105 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and informative discussion. The strongest aspects are the quantity and quality of information, while the technical level is slightly lower, reflecting the conversational format.

Reliability 8/10