Jakob Zeitler on Why LLMs Could Be Here To Stay (Even If They’re Bad) | FAI CDT

Jakob Zeitler on Why LLMs Could Be Here To Stay (Even If They’re Bad) | FAI CDT

🎙 Jakob Zeitler 👥 3K 📅 October 1, 2025 ⏱ 39 min 👁 53 📄 debate 🧭 2026-08-15
Available in: English (current) Français

Keywords

LLMproductivityevaluationstochastic parrotsAI adoption

Summary

In this interview, Jakob Zeitler discusses the trajectory of large language models (LLMs) and their potential long-term impact. He reflects on the unexpected rise of LLMs in NLP research, comparing it to past paradigm shifts like support vector machines. Zeitler argues that despite their imperfections, LLMs may become entrenched in society due to the difficulty of conducting large-scale controlled experiments to measure their true value. He explores the separation of applications based on failure tolerance, the challenge of assigning responsibility in AI-driven systems, and the philosophical question of whether LLMs possess world models or creativity. The conversation touches on the economic and resource constraints of scaling LLMs, the possibility of achieving human-level intelligence, and the risk of becoming locked into suboptimal technologies due to lack of rigorous evaluation. Zeitler emphasizes the need for better testing methods to determine where LLMs genuinely improve outcomes.

143 words

Critical Evaluation

Value of the Information & Strength of the Argument

The discussion provides valuable insights into the societal and economic factors that may sustain LLM adoption despite technical shortcomings. Zeitler’s argument that we may become ’trapped’ in suboptimal AI usage due to the impossibility of counterfactual testing is compelling and well-articulated. He balances technical skepticism with openness to possibilities, such as LLMs exhibiting primitive world models. The argumentation is logical and grounded in personal experience, though it lacks empirical evidence or references to specific studies.

Scientific Rigor, Source Quality, Title Accuracy

The conversation is rigorous in its reasoning, but it does not cite specific sources or studies. The title accurately reflects the content, which focuses on the potential persistence of LLMs despite their limitations. The discussion is more philosophical and speculative than empirical, which limits its scientific rigor but enhances its thought-provoking nature.

142 words

Title / Content Match

The title accurately reflects the central theme: the potential persistence of LLMs despite performance limitations, framed as a debate.

Quality & Reliability

7/10

The discussion is grounded in the speaker's research experience and references to known concepts (e.g., stochastic parrots, Turing test), but lacks formal citations or empirical data. The reasoning is coherent and balanced, acknowledging uncertainty.

Key Moments

Cited Sources

  • No specific sources cited in the video — The discussion references concepts like 'stochastic parrots' and 'Turing test' but does not provide direct citations.

Concurring Sources

  • No concordant sources provided — No external sources were mentioned in the video.

Dissenting Sources

  • No discordant sources provided — No external sources were mentioned in the video.

Contribution & Novelties

The interview offers a nuanced perspective on the persistence of LLMs despite their limitations, emphasizing the difficulty of rigorous evaluation. It highlights the risk of societal lock-in to suboptimal AI technologies due to the impossibility of counterfactual testing.

Pour aller plus loin :

  • Stochastic Parrots — The term is central to the debate on LLM understanding.
  • Turing Test — The discussion proposes a modified Turing test for cognitive tasks.
  • World Models — The concept of world models is discussed in relation to LLM reasoning.

84 words

Radar Profile

The radar profile shows balanced scores across information quantity, quality, technical level, and reliability, indicating a well-rounded discussion. The slightly lower reliability score reflects the lack of formal citations, while the technical level is moderate, suitable for a general audience.

Reliability 6/10