Orchestrating AI Agents to Structure the World's Product Knowledge | Shopify

Orchestrating AI Agents to Structure the World's Product Knowledge | Shopify

🎙 Kshetrajna Raghavan & Ricardo Tejedor 👥 5K 📅 November 20, 2025 ⏱ 32 min 👁 262 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

AI agentstaxonomyLLM-as-Judgeorchestrationproduction

Summary

In this talk at MLOps World 2025, Shopify engineers Kshetrajna Raghavan and Ricardo Tejedor present their system for evolving Shopify’s product taxonomy using AI agents. The taxonomy, with over 10,000 categories and 2,000 attributes, powers 30 million daily predictions. The challenge is that manual taxonomy maintenance cannot scale across all industries. They decompose the problem into four specialized agents: structure analysis (logical gaps), product analysis (data-driven gaps), merge agent (combining suggestions), and cross-link agent (handling merchant-specific categorization). Below these, an LLM-as-Judge system with nested domain-specific rules validates changes, providing states like approve, reject, approve-with-fixes, and escalate to human. They emphasize simplification, fallbacks, and audit trails for transparency. The system amplifies human experts rather than replacing them, increasing velocity while maintaining quality. Examples include adding a MagSafe compatibility attribute for smartphones. They discuss human-in-the-loop review, starting with 100% review during iteration and maintaining oversight for strategic decisions.

146 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into a real-world production AI agent system, detailing architectural patterns, challenges, and lessons learned. The argumentation is solid, based on practical experience and concrete examples. They justify design choices, such as using multiple agents for different contexts and implementing a layered judge system, with clear reasoning. The discussion of failures and the need for simplification adds credibility. The claim that the system amplifies human experts is supported by the described workflow, though quantitative metrics are limited.

90 words

Title / Content Match

The title accurately reflects the content, focusing on orchestrating AI agents for product taxonomy structuring.

Quality & Reliability

8/10

Talk by senior practitioners from Shopify with concrete production experience, clear methodology, and honest discussion of challenges. No formal peer review, but high practical credibility.

Key Moments

Cited Sources

  • MLOps World — Conference where the talk was presented.

Concurring Sources

  • Shopify Taxonomy — Official documentation of the taxonomy discussed in the talk.

Contribution & Novelties

The talk offers a detailed case study of orchestrating multiple AI agents for a complex, real-world classification task. It introduces a layered LLM-as-Judge system with domain-specific rules and a human-in-the-loop escalation mechanism. The emphasis on audit trails and transparency for trust is a valuable contribution.

Pour aller plus loin :

  • LLM-as-Judge — Overview of using LLMs for evaluation.
  • Multi-agent systems — Background on multi-agent orchestration.
  • Shopify Taxonomy — Official documentation of Shopify’s open-source taxonomy.

74 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and technical level, with slightly lower reliability due to lack of external validation. This indicates a technically rich and informative talk with practical insights, but the reliability relies on the speakers' expertise rather than peer-reviewed sources.

Reliability 8/10

💬 No comments provided.