El camino a la AGI - Guillermo Barbadillo - 2026

El camino a la AGI - Guillermo Barbadillo - 2026

🎙 Guillermo Barbadillo 👥 644 📅 January 16, 2026 ⏱ 61 min 👁 293 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

AGIARCreasoning modelsscalingintelligence

Summary

The talk, given by Guillermo Barbadillo at a Chamber of Commerce event, explores the concept of AGI and the role of the ARC benchmark in measuring progress. Barbadillo defines intelligence as the efficiency of acquiring new skills, contrasting it with mere ability. He uses the example of Deep Blue vs. Kasparov to illustrate that skill is not intelligence. He then introduces ARC, a test designed by François Chollet to measure intelligence through novel tasks that require few examples. The talk reviews the evolution of AI models on ARC: from language models showing negligible progress despite scaling, to reasoning models like OpenAI’s o3, which achieved 87% on ARC-1 in late 2024, marking a paradigm shift. The cost of evaluating o3 was $2,000,000, while the official ARC competition uses far fewer resources. In 2025, ARC-2 was released, and o3 dropped to 4%, but by the end of 2025, models reached 55% on ARC-2, and evaluation costs fell to $68. Barbadillo discusses two strategies for adapting to novelty: search (used by reasoning models) and continued learning. He concludes that current scaling approaches are insufficient for AGI and that new architectures or paradigms are needed.

191 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the distinction between skill and intelligence, and how benchmarks like ARC can measure progress toward AGI. The argumentation is coherent and well-structured, using concrete examples and data points (e.g., GPT-4’s performance on ARC, o3’s results, cost reductions). The speaker effectively argues that scaling alone has not led to AGI, and that paradigm shifts (from language models to reasoning models) are necessary. However, the argumentation is largely based on the speaker’s interpretation and lacks deep technical detail or alternative viewpoints. The talk is persuasive but not exhaustive.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor in its use of the ARC benchmark and its results, which are publicly known and verifiable. The speaker cites François Chollet and Mike Knoop, and mentions OpenAI’s o3 model, but does not provide specific citations or links. The title accurately reflects the content, which is a high-level overview of the path to AGI. The talk is not a formal scientific presentation but rather an expert opinion, which is appropriate for the context. The lack of detailed citations and the informal nature slightly reduce the rigor.

197 words

Title / Content Match

The title accurately reflects the content, which discusses the path to AGI through the lens of ARC benchmarks and recent AI progress.

Quality & Reliability

7/10

The speaker is an expert in AI, and the content is based on well-known benchmarks (ARC) and public results. However, the talk is largely opinion and interpretation, with no original data or peer-reviewed sources cited. The claims are plausible and align with known developments, but the lack of citations and the informal setting reduce the score.

Key Moments

Cited Sources

Concurring Sources

  • ARC Prize — Confirms the existence and results of the ARC competition.
  • OpenAI o3 announcement — Confirms the o3 model's performance on ARC and the reasoning paradigm.

Contribution & Novelties

The talk provides a clear and accessible explanation of the ARC benchmark and its significance in measuring progress toward AGI. It highlights the paradigm shift from language models to reasoning models and the importance of efficiency in AI. The speaker’s perspective on the limitations of scaling is valuable.

Pour aller plus loin :

  • ARC Prize — Official site with details on the competition and results.
  • François Chollet’s blog — Author’s thoughts on intelligence and AGI.
  • OpenAI o3 announcement — Official blog post about o3 and its reasoning capabilities.

88 words

Radar Profile

The radar profile shows high scores in quantity and quality of information, moderate technical level, and moderate reliability. This indicates a well-informed talk with good content but lacking deep technical detail and formal citations.

Reliability 6/10