The State of Deep Research | AI for Good New York

The State of Deep Research | AI for Good New York

🎙 Minh Trinh 👥 356 📅 October 23, 2025 ⏱ 49 min 👁 66 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

deep researchAI agentsbenchmarksreinforcement learningLLM

Summary

The talk, presented by Minh Trinh at AI for Good New York, provides a comprehensive overview of the state of deep research in AI. It begins by defining deep research agents through the perspectives of major AI providers like Claude, OpenAI, Gemini, and Perplexity, highlighting their autonomous information gathering and synthesis capabilities. The speaker then traces the evolution of these systems, from early frameworks like AutoGen and MetaGPT to more recent models such as Search-R1, ReasonRAG, and multi-agent systems like ManuSearch and DeepResearcher. The talk covers the underlying techniques, including reinforcement learning and supervised fine-tuning, and discusses hybrid approaches like EvolveSearch. A significant portion is dedicated to evaluation benchmarks, including GPQA, GAIA, Humanity’s Last Exam, WebArena, The Agent Company, and BrowseComp, explaining their purpose and challenges. The speaker also addresses risks and limitations, such as data leakage and reliability issues, and concludes with future directions, mentioning protocols like MCP and A2A. The talk is informative for a general technical audience, offering a broad survey without deep technical dives.

168 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides a valuable overview of the deep research landscape, synthesizing information from multiple sources and offering practical examples. The speaker’s argumentation is coherent, moving from definitions to technical foundations, benchmarks, and limitations. However, the presentation is largely descriptive rather than analytical, with limited critical evaluation of the models or benchmarks. The speaker occasionally interjects personal opinions, but these are not substantiated with data. The talk serves as a good introduction for those unfamiliar with the topic, but it lacks depth for experts.

Scientific Rigor, Source Quality, Title Accuracy

The speaker references several models and benchmarks, but does not provide specific citations or URLs during the talk. The only source mentioned is the speaker’s own book and website (rodeo.ai), which is not a scientific source. The talk’s rigor is moderate: it accurately describes the general concepts but does not delve into technical details or provide evidence for claims. The title accurately reflects the content, which is a state-of-the-field survey. No comments were provided for analysis.

175 words

Title / Content Match

The title accurately reflects the content, which surveys the current state of deep research AI, including models, benchmarks, and future directions.

Quality & Reliability

7/10

The talk provides a broad overview of deep research AI, covering definitions, examples, benchmarks, and limitations. The speaker demonstrates familiarity with the field, referencing multiple models and benchmarks. However, the presentation is largely descriptive and lacks deep technical detail or critical analysis. The information is generally accurate but not exhaustive, and the speaker's personal opinions are present. The talk is not peer-reviewed and is based on the speaker's expertise.

Key Moments

Markers derived by PSI from the transcript: the creator did not define chapters.

Cited Sources

Concurring Sources

  • OpenAI Deep Research — Official page describing OpenAI's deep research feature, consistent with the talk's description.
  • Anthropic Claude Research — Anthropic's research page, relevant to Claude's deep research capabilities.

Contribution & Novelties

The talk provides a comprehensive, up-to-date survey of deep research AI, synthesizing information from multiple providers and research papers. It offers a clear taxonomy of approaches, from single-agent to multi-agent systems, and discusses key benchmarks and their limitations. The speaker’s perspective as a practitioner adds practical insights, though the talk is more descriptive than analytical.

Pour aller plus loin :

100 words

Radar Profile

The radar profile shows high scores in quantity of information and technical level, indicating a dense and informative talk. However, the quality of information and global reliability are slightly lower, reflecting the lack of detailed citations and the speaker's subjective viewpoint. The overall balance suggests a useful overview for those seeking a broad understanding of deep research AI.

Reliability 7/10