How to train your LLM | AI Talk 54

How to train your LLM | AI Talk 54

🎙 Wasner + Steinschaden - Der KI-Podcast 👥 242 📅 November 26, 2025 ⏱ 48 min 👁 90 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

fine-tuningLLMMistralon-deviceAI Act

Summary

In this episode of AI Talk, hosts Jakob Steinschaden and Clemens Wasner interview Matthias Neumayr and Dima Rubanov, founders of Oscar Stories and HeyQQ. They discuss their journey of building a custom LLM for generating child-friendly content. The conversation covers the limitations of prompting, the decision to fine-tune an open-source model, and the detailed process of data collection, cleaning, and model selection. They chose Mistral 7B for its superior German language performance and fine-tuned it using LoRA on a MacBook M3, a process that took nearly 24 hours. They emphasize the importance of benchmarking and the challenges of hosting and inference costs. The discussion also touches on on-device AI as a future direction for privacy, the TÜV Trusted AI certification they obtained, and the practical implications of the EU AI Act. The episode provides practical insights for startups considering fine-tuning their own models.

143 words

Critical Evaluation

Value of the Information & Strength of the Argument

The value of the information lies in its practical, hands-on perspective on fine-tuning an LLM for a specific use case. The guests share concrete details about the process, including data cleaning, model selection, and benchmarking, which is valuable for practitioners. The argumentation is based on their own experience rather than rigorous scientific evidence, but it is coherent and grounded in real-world challenges. They also provide a balanced view by discussing the limitations of fine-tuning and the importance of trying simpler methods first.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is moderate. The guests mention specific models and techniques but do not provide detailed quantitative results or comparisons. They reference their own benchmarking but do not share the methodology or data. The sources cited in the description include their product website and a Financial Times article, but these are not directly used to support the technical claims. The title accurately reflects the content, which is a practical guide to fine-tuning an LLM.

172 words

Title / Content Match

The title 'How to train your LLM' accurately reflects the core topic of the episode, which focuses on the practical steps and considerations for fine-tuning an LLM for a startup use case.

Quality & Reliability

7/10

The hosts and guests are practitioners with hands-on experience in fine-tuning LLMs for a specific use case. They provide concrete details about the process, challenges, and model selection. However, the discussion is largely anecdotal and lacks rigorous scientific methodology or external validation. The video is a podcast conversation, not a peer-reviewed study.

Key Moments

Cited Sources

  • Oscar Stories — The product website for the app discussed in the episode.
  • FT - Insurers retreat from AI cover as risk of multibillion-dollar claims mounts — Mentioned in the description, likely related to AI risk and insurance.

Concurring Sources

  • Oscar Stories — The product website for the app discussed in the episode.

Contribution & Novelties

The episode provides a practical, first-hand account of fine-tuning an LLM for a niche application (child-friendly content) in a startup context. It highlights the challenges of data collection, model selection, and benchmarking, and offers insights into the trade-offs between using large API-based models and self-hosted smaller models. The discussion on on-device AI and privacy is particularly relevant for consumer applications.

Pour aller plus loin :

122 words

Radar Profile

The radar profile shows a balanced performance across all dimensions, with slightly higher scores in information quantity and quality, and lower in technical depth and reliability. This indicates a solid but not deeply technical discussion, suitable for a general audience interested in practical AI applications.

Reliability 6/10

💬 No comments were provided for analysis.