A Practical Field Guide to Optimizing Cost, Speed & Accuracy of LLMs | Niels Bantilan, Union.ai

A Practical Field Guide to Optimizing Cost, Speed & Accuracy of LLMs | Niels Bantilan, Union.ai

🎙 Niels Bantilan 👥 5K 📅 October 20, 2025 ⏱ 33 min 👁 46 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

SLMLLMoptimizationorchestrationcost

Summary

Niels Bantilan presents a practical field guide for optimizing cost, speed, and accuracy of LLMs in domain-specific agents, focusing on the adoption of Small Language Models (SLMs). He begins by referencing Nvidia’s position paper arguing that SLMs are the future of agentic AI, citing their sufficiency, operational suitability, and economic benefits. He acknowledges barriers such as upfront investment and maintenance. He then outlines a pain-driven development journey from initial prototype to scaling issues, highlighting the need for evaluation, monitoring, and refactoring. The core strategy involves identifying leverage points to swap LLMs for SLMs, maintaining performance through AI unit tests and synthetic datasets, optimizing experimentation speed via prompt optimization and fine-tuning, managing costs with heterogeneous compute, and optimizing inference speed through container reuse and auto-scaling. He uses a SQL agent as a case study, demonstrating how resource-aware AI-native orchestration (specifically Flyte 2) can facilitate these optimizations. He concludes that while SLMs offer significant savings, they require investment in tools and practices, and emphasizes that agent code is trivial compared to the supporting infrastructure.

172 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable, actionable insights for AI/ML teams considering SLM adoption. The argumentation is coherent and grounded in practical experience, using a concrete case study (SQL agent) to illustrate strategies. The speaker effectively addresses cost, speed, and accuracy trade-offs, and emphasizes the importance of orchestration and evaluation. However, the talk is largely based on anecdotal evidence and the speaker’s own experience, lacking rigorous empirical data or comparative studies. The argumentation is persuasive but not deeply technical, and some claims (e.g., 100x savings) are presented as optimistic estimates without detailed calculations.

Scientific Rigor, Source Quality, Title Accuracy

The talk references two key sources: Nvidia’s position paper on SLMs and the ‘Hidden Technical Debt in Machine Learning Systems’ paper. These are relevant and credible, but the talk does not provide detailed citations or URLs. The title accurately reflects the content, focusing on practical optimization strategies. The talk is well-structured and the speaker is knowledgeable, but the lack of explicit citations and the reliance on personal experience somewhat reduce the scientific rigor. The description includes a link to MLOps World, but no direct links to the cited papers.

195 words

Title / Content Match

The title accurately reflects the content: a practical guide to optimizing cost, speed, and accuracy of LLMs, with a focus on SLMs.

Quality & Reliability

7/10

The talk is an expert opinion based on practical experience, referencing the Nvidia position paper and the hidden technical debt paper. It provides concrete strategies and examples, but lacks detailed citations and empirical validation.

Key Moments

Cited Sources

  • MLOps World — Conference website for the event where the talk was given.

Concurring Sources

Contribution & Novelties

The talk provides a practical, step-by-step framework for adopting SLMs in production, emphasizing the importance of resource-aware orchestration. It offers concrete strategies for identifying leverage points, maintaining performance, and optimizing cost and speed. The use of a SQL agent case study makes the concepts tangible. The talk also highlights the hidden technical depth in agent systems, drawing an analogy to the ‘Hidden Technical Debt’ paper.

Pour aller plus loin :

  • Small language models (SLMs) — Overview of SLMs and their characteristics.
  • Prompt engineering — Techniques for optimizing prompts, relevant to the talk’s prompt optimization strategy.
  • LLM-as-a-judge — Reference for using LLMs as evaluators, as mentioned in the talk.
  • Flyte — The orchestration platform mentioned in the talk, for building AI-native workflows.

121 words

Radar Profile

The radar profile shows high scores in information quantity and quality, reflecting the talk's practical depth. The technical level is moderate, suitable for a broad audience. The global reliability is good, but the lack of detailed citations slightly lowers the score.

Reliability 7/10