29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy Berman

29.4% ARC-AGI-2 🤯 (TOP SCORE!) - Jeremy Berman

🎙 Jeremy Berman 👥 218K 📅 September 27, 2025 ⏱ 68 min 👁 18K 📄 interview 🧭 2026-08-15
Available in: English (current) Français

Keywords

ARC-AGIevolutionnatural languagereasoningRL

Summary

In this interview, Jeremy Berman discusses his recent top score on the ARC-AGI-2 leaderboard using an evolutionary approach that generates natural language descriptions of transformation rules instead of Python code. He explains the shift from code to language, highlighting the expressiveness of English for describing ARC tasks. The conversation covers the role of reinforcement learning in improving reasoning, the distinction between memorized and deduced knowledge, and the challenges of catastrophic forgetting and continual learning. Berman argues that the meta-skill of reasoning is key to AGI, and discusses the potential for composable models and active inference. The discussion also touches on the nature of intelligence, the limitations of current LLMs, and future directions for AI research.

115 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable insights into a novel approach to ARC-AGI, with Berman clearly explaining his methodology and the rationale behind it. The argumentation is solid, grounded in practical results and references to relevant research. Berman’s distinction between memorized and deduced knowledge, and his emphasis on the meta-skill of reasoning, are thought-provoking. The discussion is balanced, with the host challenging Berman’s views, leading to a nuanced exploration of the limits and potential of current AI.

Scientific Rigor, Source Quality, Title Accuracy

The video references several credible sources, including Berman’s own blog posts, the ARC-AGI benchmark, and academic papers. The discussion is technically rigorous, with specific references to concepts like the ‘LLM biology’ paper and the ‘Fractured Entangled Representation Hypothesis’. The title accurately reflects the content, focusing on Berman’s achievement and the subsequent discussion. No significant discrepancies between title and content were noted.

151 words

Title / Content Match

The title accurately reflects the main topic: Jeremy Berman's top score on ARC-AGI-2 and the discussion around it.

Quality & Reliability

8/10

High-quality technical discussion with a leading researcher, grounded in specific references and personal experience, though claims are forward-looking and not peer-reviewed.

Key Moments

Cited Sources

Concurring Sources

Dissenting Sources

  • On the Measure of Intelligence — Chollet's definition of intelligence as skill-acquisition efficiency might be interpreted as conflicting with Berman's emphasis on reasoning as a meta-skill, though both emphasize generalization.

External References

Contribution & Novelties

The video provides a unique insight into a novel approach to ARC-AGI that uses natural language descriptions instead of code, highlighting the expressiveness of language for abstract reasoning tasks. Berman’s distinction between memorized and deduced knowledge, and his emphasis on the meta-skill of reasoning, offer a fresh perspective on AGI research. The discussion also touches on the challenges of continual learning and the potential for composable models, which are underexplored areas.

Pour aller plus loin :

125 words

Radar Profile

The radar profile shows high scores in information quantity and quality, with a moderate technical level. The reliability is also high, indicating a well-grounded discussion. The profile suggests a content that is informative and credible, but not overly technical for a general audience.

Reliability 8/10

💬 Positif. Les commentaires sont largement favorables, saluant la profondeur de la discussion et les références, avec quelques critiques constructives sur les limites du raisonnement des LLM.