Google’s Gemini 3: AI agents, reasoning and search mode

Google’s Gemini 3: AI agents, reasoning and search mode

🎙 IBM Technology 👥 1.8M 📅 November 21, 2025 ⏱ 46 min 👁 28K 📄 expert opinion 🧭 2026-08-06
Available in: English (current) Français

Keywords

Gemini 3AI agentshallucinationbenchmarksenterprise AI

Summary

In this episode of Mixture of Experts, host Tim Hwang and IBM experts Marina Danilevsky, Gabe Goodhart, and Merve Unuvar discuss the recent release of Google’s Gemini 3 model. They analyze its benchmark performance, noting impressive gains on difficult evals like Humanity’s Last Exam and ARC-AGI, but also highlight persistent issues with hallucination and a tendency to provide answers rather than admit uncertainty. The conversation shifts to Google’s new agentic IDE, Antigravity, which the panel sees as a key differentiator, enabling users to manage fleets of delegate agents. Merve shares her hands-on experience building a workout dashboard, noting both the model’s strengths and its occasional errors. The discussion then covers IBM’s CUGA agent framework, designed for enterprise-ready generalist agents, and the broader trend toward specialized agents rather than a single all-purpose model. The panel also debates OpenAI’s GDPval benchmark on AI’s economic impact, questioning its assumptions about job automation. Finally, they discuss Anthropic’s disruption of an AI-led cyberattack, raising concerns about AI agents’ potential for malicious use. Throughout, the experts emphasize the importance of moving beyond benchmarks to real-world applications and the need for a suite of specialized models rather than one model to rule them all.

197 words

Critical Evaluation

The podcast provides a balanced and insightful discussion of recent AI developments, particularly the release of Gemini 3. The panelists, all with deep technical expertise, offer a mix of hands-on experience and analytical commentary. Marina’s observation about Gemini 3’s persistent hallucination despite benchmark improvements is a valuable critique, highlighting the gap between benchmark performance and real-world reliability. Gabe’s perspective on ecosystem moats and the differentiation through Antigravity is well-argued, emphasizing the shift from raw model capability to integrated agentic workflows. Merve’s practical test of the model, building a dashboard, adds a concrete dimension to the discussion, though her anecdote about the model’s error in assuming she wanted to ‘grow’ illustrates the limitations of current models in understanding context. The discussion on IBM’s CUGA framework provides insight into enterprise agent development, though it is brief and could benefit from more technical detail. The debate on OpenAI’s GDPval benchmark raises important questions about the economic impact of AI, but the panel’s skepticism is not deeply substantiated with data. The segment on Anthropic’s cyberattack disruption is timely and highlights the dual-use nature of AI, but again, the analysis is more speculative than evidence-based. Overall, the podcast is informative and engaging, offering expert opinions rather than rigorous scientific analysis. The lack of citations to specific sources within the discussion is a minor weakness, but the panel’s credibility and the range of topics covered make it a valuable resource for those following AI trends. The adéquation between title and content is good, though the title focuses on Gemini 3 while the episode covers multiple topics. The public comments, if any, are not provided, so no analysis of audience reception is possible.

276 words

Title / Content Match

The title accurately reflects the main focus on Gemini 3, but the video also covers other AI topics, making it slightly broader than the title suggests.

Quality & Reliability

7/10

The podcast features expert opinions from IBM researchers and architects, discussing recent AI developments with a mix of personal experience and analysis. While not a formal study, the panel provides informed perspectives, but the lack of detailed technical verification and reliance on anecdotal evidence slightly reduce the score.

Chapters

Cited Sources

Concurring Sources

  • Gemini 3 announcement — Aligns with the panel's discussion of Gemini 3's capabilities and benchmarks.
  • ARC-AGI benchmark — Supports the discussion on Gemini 3's performance on this benchmark.

Dissenting Sources

  • OpenAI GDPval

Contribution & Novelties

The podcast offers a timely expert analysis of Gemini 3’s release, providing hands-on impressions and contextualizing its benchmark performance within the broader AI ecosystem. The discussion on Antigravity as a differentiator and the shift toward agentic workflows adds value beyond typical news coverage. The debate on GDPval and AI’s economic impact introduces critical perspectives on AI’s societal implications.

Pour aller plus loin :

  • Gemini 3 official page — Official announcement and technical details.
  • ARC-AGI benchmark — The benchmark referenced for measuring AI’s reasoning capabilities.
  • Humanity’s Last Exam — The benchmark mentioned for evaluating advanced AI knowledge.
  • IBM CUGA framework — More details on IBM’s agent framework discussed in the episode.
  • OpenAI GDPval — The benchmark on AI’s economic impact debated in the episode.

123 words

Radar Profile

The radar profile shows a balanced performance across all dimensions, with slightly higher scores in information quantity and reliability, reflecting the podcast's informative yet opinion-based nature. The lower technical depth score indicates that while the discussion is expert-level, it remains accessible to a broader audience.

Reliability 7/10