RAG: The 2025 Best-Practice Stack, Prototype to Production

RAG: The 2025 Best-Practice Stack, Prototype to Production

🎙 Toronto Machine Learning Society (TMLS) 👥 5K 📅 September 25, 2025 ⏱ 172 min 👁 394 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

RAGLLMLangGraphQdrantRAGAS

Summary

This workshop, presented by Greg Loughnane and Chris Alexiuk of AI Makerspace, provides a comprehensive guide to building production-ready RAG (Retrieval-Augmented Generation) applications. The speakers share their recommended ‘best-practice stack’ for 2025, which includes LangGraph for orchestration, Qdrant as the vector database, Cohere’s Rerank for improved retrieval, RAGAS for evaluation, Together AI for model serving, and Llama 3.3 70B as the LLM. They emphasize the importance of moving from simple prototypes to production-grade systems, outlining a five-phase approach that includes on-prem demos, data preparation, beta testing, and scaling. The session includes live coding demonstrations, practical advice on monitoring and evaluation, and a discussion of the trade-offs between cloud and on-premise deployments. The speakers also address common enterprise pain points and highlight the need for metrics-driven development. The workshop concludes with a Q&A session, offering attendees the opportunity to engage with the presenters and delve deeper into specific topics.

148 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video offers significant practical value by presenting a concrete, actionable stack for RAG development, based on the speakers’ extensive experience in AI engineering education and consulting. The argumentation is solid, as they justify each tool choice with clear reasoning, such as LangGraph’s flexibility for complex systems, Qdrant’s scalability, and RAGAS’s credibility in evaluation. They also provide a realistic phased approach to production, acknowledging the challenges of enterprise adoption. However, the content is primarily opinion-based, lacking formal citations or comparative benchmarks, which limits its scientific rigor.

95 words

Title / Content Match

The title accurately reflects the content, which focuses on the best-practice RAG stack for 2025 and the journey from prototype to production.

Quality & Reliability

7/10

The video presents a well-structured, practical overview of a recommended RAG stack, drawing on the speakers' extensive industry experience. Claims are substantiated with reasoning and comparisons, but the content is largely opinion-based and lacks formal citations or empirical validation.

Key Moments

Cited Sources

Concurring Sources

  • A16Z LLM Stack — Referenced as the foundational architecture for LLM applications.

Contribution & Novelties

The video provides a clear, up-to-date recommendation for a production-ready RAG stack, synthesizing current best practices into a single actionable framework. It emphasizes the importance of evaluation and monitoring, and offers a phased approach to enterprise adoption. The practical coding demonstrations and real-world insights from industry practitioners add significant value.

Pour aller plus loin :

111 words

Radar Profile

The radar profile shows high scores in information quantity and technical level, indicating a content-rich and technically detailed presentation. The lower scores in information quality and reliability reflect the opinion-based nature of the recommendations, which are not backed by formal research.

Reliability 6/10

💬 No comments were provided for analysis.