From Vectors to Agents: Managing RAG in an Agentic World

From Vectors to Agents: Managing RAG in an Agentic World

🎙 Rajiv Shah 👥 5K 📅 October 23, 2025 ⏱ 85 min 👁 153 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

RAGBM25Embedding ModelsAgentic SearchContext Engineering

Summary

Rajiv Shah, Chief Evangelist at Contextual AI, presents a workshop on managing RAG systems in an agentic world. He begins by highlighting the evolution from simple keyword search to semantic embeddings and now agentic reasoning. He emphasizes the importance of understanding use cases and trade-offs such as latency, cost, problem complexity, and cost of mistakes. The talk focuses on retrieval, covering three main approaches: BM25, language models (embeddings), and agentic search. BM25 is presented as a strong baseline using inverted indexes, with limitations in handling synonyms and abbreviations. Language models, particularly contextualized embeddings, are discussed with references to the MTEB leaderboard for model selection. He introduces the concept of ‘Speedy Retrieval’ (500ms), ‘Accuracy-Optimized RAG’ (10s), and ‘Exhaustive Agentic Search’ for complex reasoning. He provides decision frameworks and practical advice on context engineering and window management. The talk concludes with a discussion on when to use speed-first retrieval versus agentic search, emphasizing that ‘good enough’ retrieval often beats ‘perfect’ agentic reasoning in production. He also touches on cost, latency, and complexity comparisons, and offers a framework for designing scalable RAG architectures.

180 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable practical insights into RAG system design, drawing on the speaker’s extensive experience. It offers a clear framework for choosing between different retrieval approaches based on latency, cost, and complexity. The argumentation is solid, with concrete examples and references to benchmarks like MTEB. However, it lacks formal citations and rigorous empirical validation, relying on anecdotal evidence and personal experience. The speaker effectively communicates trade-offs and encourages audience questions, enhancing the value of the session.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor through references to established benchmarks (MTEB, BEIR) and leaderboards, and the speaker’s background at Hugging Face adds credibility. However, specific sources are not cited in the talk, and the description only provides a link to the conference. The title accurately reflects the content, which transitions from basic vector-based RAG to agentic approaches. The talk is well-structured and practical, but the lack of formal citations and reliance on personal experience slightly reduce its scientific rigor.

171 words

Title / Content Match

The title accurately reflects the content, which transitions from basic vector-based RAG to agentic approaches, focusing on managing RAG in production.

Quality & Reliability

7/10

The talk provides practical, experience-based guidance on RAG architectures, with references to benchmarks and leaderboards. However, it lacks formal citations and rigorous empirical validation, relying on anecdotal evidence and personal experience.

Key Moments

Cited Sources

  • MLOps World — Conference where the talk was recorded.

Concurring Sources

Contribution & Novelties

The talk provides a practical framework for choosing between different RAG architectures based on latency, cost, and complexity. It emphasizes the importance of context engineering and offers decision frameworks for production-ready systems. The speaker’s experience at Hugging Face and Contextual AI adds credibility.

Pour aller plus loin :

85 words

Radar Profile

The radar profile shows high scores in quantity of information and technical level, indicating a content-rich talk. Quality of information and global reliability are moderate, reflecting the practical but non-rigorous nature. The talk is strong in providing actionable insights but lacks formal citations.

Reliability 6/10