Observability Panel | Galileo, DraftKings, Target Corporation, PIMCO | MLOps World 2025

Observability Panel | Galileo, DraftKings, Target Corporation, PIMCO | MLOps World 2025

🎙 Toronto Machine Learning Society (TMLS) 👥 5K 📅 October 23, 2025 ⏱ 32 min 👁 112 📄 panel discussion 🧭 2026-08-15
Available in: English (current) Français

Keywords

LLM observabilitymonitoringevaluationgroundinggolden dataset

Summary

This panel discussion, recorded at MLOps World 2025, brings together experts from Galileo, DraftKings, Target, and PIMCO to discuss the evolving landscape of LLM observability. The speakers emphasize that traditional ML observability is insufficient for LLM-based systems due to their non-deterministic and black-box nature. Key topics include the importance of end-to-end tracing, logging, evaluation, and human-in-the-loop feedback. Atin Sanyal from Galileo discusses the challenges of scaling evaluations, introducing their low-latency evaluation model ‘Luna’ as an alternative to LLM judges. Naresh Kumar Batthula from DraftKings highlights the need for context lineage and a golden dataset as a starting point for testing. Bali Varadarajan from Target stresses the importance of grounding and cost monitoring for building trust. Naveen from realtor.com (though listed as PIMCO in the description) discusses unifying observability signals to avoid alert fatigue and emphasizes functional and outcome observability. The panel concludes that observability is critical for production AI, but must be tailored to business use cases and include both technical and business metrics.

164 words

Critical Evaluation

Value of the Information & Strength of the Argument

The panel provides valuable insights from practitioners with hands-on experience in deploying LLM systems at scale. The discussion covers a range of perspectives, from technical implementation (tracing, logging, evaluation) to business considerations (cost, trust, user feedback). The argumentation is largely based on anecdotal evidence and personal experience, which is appropriate for a panel format. However, the lack of concrete data or formal studies limits the scientific rigor. The speakers do not always provide detailed explanations, but the overall discussion is coherent and addresses key challenges in LLM observability.

Scientific Rigor, Source Quality, Title Accuracy

The panel is composed of industry experts, but no formal sources are cited during the discussion. The only external reference is the MLOps World website, which is not a scientific source. The title accurately reflects the content, as it is indeed an observability panel with the mentioned companies. The discussion is practical and grounded in real-world experience, but the lack of citations and reliance on anecdotal evidence means the scientific rigor is moderate. The content is more of an expert opinion than a peer-reviewed study.

188 words

Title / Content Match

The title accurately reflects the content: a panel discussion on observability with representatives from the mentioned companies.

Quality & Reliability

7/10

Panel of industry practitioners from major companies discussing LLM observability. Provides practical insights and real-world examples, but lacks formal citations and is based on anecdotal experience.

Key Moments

Cited Sources

  • MLOps World — Conference website where the panel was recorded.

Concurring Sources

  • MLOps World — Conference website confirming the event and speakers.

Contribution & Novelties

The panel provides a practical overview of LLM observability from multiple enterprise perspectives, highlighting the shift from traditional ML monitoring to more holistic approaches. It emphasizes the importance of tracing, logging, evaluation, and human feedback, and introduces innovative solutions like low-latency evaluation models (Luna) to address scalability challenges. The discussion also underscores the need for context lineage and golden datasets to ensure reliability.

Pour aller plus loin :

102 words

Radar Profile

The radar profile shows high scores in information quantity and quality, reflecting the panel's rich content and practical insights. The technical level is moderate, suitable for a broad audience, while reliability is decent but limited by the lack of formal citations. Overall, the panel offers valuable guidance for practitioners.

Reliability 7/10

💬 No comments were provided for analysis.