Estonian language in large language models

Estonian language in large language models

🎙 Kairit Sirts 👥 1K 📅 September 26, 2025 ⏱ 26 min 👁 137 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

EstonianLLMevaluationopen modelslanguage resources

Summary

Kairit Sirts, Associate Professor of Language Technology at the University of Tartu, presents an overview of the integration of Estonian into large language models (LLMs). She traces the evolution from early English-only models like GPT and BERT to multilingual models and the recent open-weight Llama family. She highlights that Estonian is not yet at the same level as English due to data scarcity. The talk addresses two main questions: how to evaluate the quality of Estonian in LLMs and whether local efforts are needed to improve it. She describes four evaluation methods: subjective impression, automatic benchmarks, crowd-sourced evaluation via Chatbot Arena (including the Estonian variant Tehisaru), and LLM-as-a-judge. She argues for local involvement in model development to leverage local data, build competence, and handle sensitive data, while dismissing the idea of training an Estonian GPT from scratch. The project SLLM focuses on data collection, continued training of open models, and evaluation. The talk concludes with a Q&A session discussing technical details and licensing.

163 words

Critical Evaluation

Value of the Information & Strength of the Argument

The presentation provides valuable insights into the state of Estonian in LLMs, backed by practical experience from the SLLM project. The argumentation is coherent and well-structured, moving from historical context to evaluation methods and then to the rationale for local development. The speaker effectively balances technical details with broader considerations, such as data rights and societal benefits. The use of concrete examples, like the Chatbot Arena and the prototype model, strengthens the credibility of the claims. However, some points, such as the effectiveness of continued training, are based on preliminary results and could benefit from more extensive evidence.

Scientific Rigor, Source Quality, Title Accuracy

The presentation demonstrates scientific rigor through the systematic overview of evaluation methods and the discussion of specific models and projects. The speaker references well-known models (GPT, BERT, Llama) and initiatives (Chatbot Arena), but does not provide formal citations or links to specific papers. The title accurately reflects the content, focusing on the status and development of Estonian in LLMs. The talk is based on the speaker’s expertise and ongoing research, which adds credibility, but it is not a peer-reviewed source. The Q&A session further clarifies technical aspects, showing transparency and depth of knowledge.

206 words

Title / Content Match

The title accurately reflects the content, focusing on the status and development of Estonian in large language models.

Quality & Reliability

8/10

The presentation is given by a qualified associate professor in language technology, based on ongoing research and practical experience. It provides a balanced view of the challenges and approaches, with references to specific models and projects. However, it is an expert opinion rather than a peer-reviewed study, and some claims are not fully detailed.

Key Moments

Cited Sources

  • Chatbot Arena — Mentioned as a platform for crowd-sourced evaluation of LLMs.
  • Tehisaru — Estonian version of Chatbot Arena developed by the speaker's team.
  • Llama 2 — Open-weight model family from Meta, used as a base for the SLLM project.

Concurring Sources

  • Estonian Language Technology Programme — Funding source for the SLLM project, supporting language technology development in Estonia.

Contribution & Novelties

The presentation offers a unique perspective on the challenges and opportunities of integrating a low-resource language like Estonian into LLMs. It provides practical insights from an ongoing national project, including the development of an Estonian evaluation platform (Tehisaru) and the continued training of open models. The talk emphasizes the importance of local competence and data sovereignty, which is a valuable contribution to the discourse on AI and language preservation.

Pour aller plus loin :

  • Large language model — Overview of LLMs and their development.
  • Transfer learning — Key concept for adapting pre-trained models to new languages.
  • Low-resource languages — Challenges and strategies for NLP in languages with limited data.

109 words

Radar Profile

The radar profile shows high scores in quality, technical level, and reliability, reflecting the speaker's expertise and the well-structured content. The quantity of information is moderate, as the talk is concise and focused. The overall balance indicates a reliable and informative presentation.

Reliability 8/10

💬 No comments were provided for analysis.