Architecting a Deep Research System | Suhas Pai, Hudson Labs

Architecting a Deep Research System | Suhas Pai, Hudson Labs

🎙 Suhas Pai 👥 5K 📅 October 20, 2025 ⏱ 29 min 👁 142 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

deep researchretrievalreasoningevaluationagentic AI

Summary

Suhas Pai, CTO of Hudson Labs, presents a talk on architecting deep research systems, drawing from his experience building one for the financial domain. He defines deep research as a system that takes arbitrary retrieval and reasoning hops to resolve an information need in detail, emphasizing the interplay between retrieval and reasoning. He illustrates the potential with an anecdote where the system deduced a supply chain issue from two disparate documents. He then breaks down the core components: retrieval system (agentic retrieval, query generation, top-down vs bottom-up), reasoning system (using models like OpenAI o1 or DeepSeek R1), report generation system (long-context LLM), depth and breadth balancer, and context orchestrator. He discusses tradeoffs like internal vs external decision-making and context engineering. For evaluation, he proposes inventing scenarios with emergent knowledge, using Wikipedia as a test set, and evaluating at component level with metrics like factuality, citability, and non-repetitiveness. He addresses questions on data source coverage and hallucinations, recommending building from scratch and incorporating verification steps. The talk concludes with advice for beginners to start with a simple RAG-based system.

178 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable practical insights into building deep research systems, based on the speaker’s direct experience. The argumentation is coherent and structured, moving from definition to components to evaluation. The anecdote about the supply chain deduction effectively illustrates the potential value. The discussion of tradeoffs (e.g., internal vs external decision-making, depth vs breadth) is nuanced and useful. However, the talk lacks empirical evidence or benchmarks to support claims, and some suggestions (like inventing scenarios) are presented without detailed methodology. The Q&A section adds practical advice but remains high-level.

Scientific Rigor, Source Quality, Title Accuracy

The talk is an expert opinion piece with no formal citations or references to external sources. The only link provided is to the MLOps World conference, which is not a source of technical content. The speaker mentions his book but does not provide a URL. The title accurately reflects the content, which is focused on architectural considerations. The lack of sources reduces the scientific rigor, but the practical experience lends credibility. The talk does not reference specific research papers or benchmarks, so the quality of sources is low.

192 words

Title / Content Match

The title accurately reflects the content, which focuses on architectural components and tradeoffs of deep research systems.

Quality & Reliability

7/10

The speaker is CTO of Hudson Labs with practical experience building deep research systems. The talk provides concrete architectural insights and evaluation strategies, but lacks formal citations and empirical validation. Claims are plausible and grounded in industry experience.

Key Moments

Cited Sources

  • MLOps World — Conference website where the talk was presented

Concurring Sources

  • OpenAI Deep Research — OpenAI's deep research system, mentioned as the first of its kind.
  • Google Gemini Deep Research — Google's deep research feature, mentioned in the talk.

Contribution & Novelties

The talk provides a practical, component-based framework for architecting deep research systems, emphasizing the importance of depth/breadth balancing and context engineering. It offers novel evaluation strategies, such as inventing scenarios with emergent knowledge and using Wikipedia as a test set. The speaker’s experience in the financial domain adds a unique perspective.

Pour aller plus loin :

93 words

Radar Profile

The radar profile shows high scores in information quantity and technical level, but lower in reliability due to lack of citations. The overall quality is good, with a strong practical focus.

Reliability 6/10