LLMOps Infrastructure for Production-Grade Agentic RAG Applications with Union.ai

LLMOps Infrastructure for Production-Grade Agentic RAG Applications with Union.ai

🎙 Niels Bantilan 👥 5K 📅 September 28, 2025 ⏱ 86 min 👁 203 📄 tutorial 🧭 2026-08-15
Available in: English (current) Français

Keywords

RAGLLMOpsproductionevaluationhyperparameter optimization

Summary

This workshop, presented by Niels Bantilan, Chief ML Engineer at Union.ai, focuses on the operational and infrastructure challenges of moving RAG (Retrieval-Augmented Generation) applications from prototype to production. The speaker emphasizes that unlike traditional ML models, RAG systems are compound AI systems requiring careful management of components like vector stores, embedding models, and prompts. The session covers four main sections: building a basic RAG pipeline, creating an evaluation dataset using LLM-as-a-judge, applying hyperparameter optimization (HPO) to RAG configurations, and orchestrating the entire pipeline using Union.ai’s serverless platform. Attendees are guided through a hands-on notebook to implement these concepts, with a focus on systematic improvement and best software engineering practices. The talk highlights the importance of treating RAG pipelines as tunable systems and using evaluation metrics to guide optimization, rather than relying on ad-hoc approaches.

134 words

Critical Evaluation

Value of the Information & Strength of the Argument

The value of this workshop lies in its practical, systems-oriented approach to RAG deployment. The speaker effectively argues that RAG applications require continuous iteration and a structured methodology similar to traditional ML, but with additional complexity. He introduces the concept of treating the entire RAG pipeline as a black-box function for hyperparameter optimization, which is a valuable mental model. The argumentation is solid, grounded in real-world experience, and supported by concrete examples. The speaker also addresses the challenge of evaluation in LLM-based systems, advocating for LLM-as-a-judge as a pragmatic solution despite its limitations. The workshop provides actionable insights and a clear framework for practitioners.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is good, with the speaker drawing on his experience as a core maintainer of Flyte and creator of Pandera. The sources cited are primarily the tools and platforms used (Union.ai, Flyte, Pandera) and the pandas documentation. The title accurately reflects the content, which is focused on LLMOps infrastructure for RAG applications. The workshop is well-structured and the technical details are accurate. However, the content is somewhat promotional for Union.ai, and the evaluation methods (LLM-as-a-judge) are acknowledged as having limitations. No comments were provided for analysis.

207 words

Title / Content Match

The title accurately reflects the content, which focuses on LLMOps infrastructure for production-grade RAG applications, with a practical demonstration using Union.ai.

Quality & Reliability

8/10

The speaker is a chief ML engineer with deep expertise in MLOps, and the content is practical and well-structured. However, it is a workshop with a promotional component for Union.ai, and the evaluation methods rely on LLM-as-a-judge, which has known limitations.

Key Moments

Cited Sources

  • Union.ai — The platform used for the workshop, providing serverless compute for ML workflows.
  • Flyte — Open-source workflow orchestration tool, core maintainer of which is the speaker.
  • Pandera — Data validation tool for dataframes, created by the speaker.
  • Pandas documentation — Used as the knowledge base for the RAG application.

Concurring Sources

Dissenting Sources

Contribution & Novelties

The workshop provides a practical, hands-on approach to LLMOps for RAG applications, emphasizing the importance of treating RAG pipelines as tunable systems. It introduces a systematic methodology for evaluation and hyperparameter optimization, which is often overlooked in favor of more complex agentic approaches. The use of Union.ai as a platform demonstrates how to orchestrate the entire pipeline, from data ingestion to evaluation and optimization.

Pour aller plus loin :

104 words

Radar Profile

The radar profile shows high scores in information quantity and quality, reflecting the workshop's comprehensive coverage. The technical level is high, indicating advanced content. The reliability score is slightly lower due to the promotional aspect and reliance on LLM-as-a-judge.

Reliability 7/10