Big Llm architecture comparison - part 2

Big Llm architecture comparison - part 2

🎙 West Coast Machine Learning 👥 3K 📅 October 8, 2025 ⏱ 82 min 👁 88 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

Kimi K2GPT-OSSGrok-2.5GLM-4.5Qwen-3

Summary

This meetup video is the second part of a discussion on Sebastian Raschka’s article ‘The Big LLM Architecture Comparison’. The participants, likely ML engineers and researchers, analyze several recent large language models, focusing on their architectural choices. They start with Kimi K2, noting its use of the Muon optimizer and its similarity to DeepSeek V3, with a large number of experts (384) but only 8 active. They then discuss GPT-OSS, expressing skepticism about OpenAI’s motivations and highlighting its attention bias mechanism as an alternative to attention sinks. Next, they cover Grok-2.5, which they find fairly standard, and GLM-4.5, which is a large MoE model. The discussion includes comparisons of width vs. depth, expert counts, and shared experts. The tone is informal and technical, with participants sharing insights and asking clarifying questions.

131 words

Critical Evaluation

Value of the Information & Strength of the Argument

The value of the information lies in the practical insights from practitioners who are actively working with these models. They provide context on trade-offs, such as the business incentives behind architectural decisions (e.g., OpenAI’s GPT-OSS) and the practical implications of expert collocation for inference. The argumentation is largely based on personal experience and reasoning, with references to the cited article and papers. While not rigorously scientific, the discussion is informed and highlights nuances that might be missed in a formal paper.

Scientific Rigor, Source Quality, Title Accuracy

The discussion is grounded in Sebastian Raschka’s article, which is a credible source. The participants reference specific details from the paper and their own knowledge. However, they often speculate or admit uncertainty, which reduces the rigor. The title accurately reflects the content, as it is a continuation of a comparison of LLM architectures. No comments were provided for analysis.

155 words

Title / Content Match

The title accurately reflects the content, which is a continuation of a comparative analysis of large language model architectures.

Quality & Reliability

7/10

Discussion among ML practitioners, grounded in a specific article by Sebastian Raschka, but with informal reasoning and some uncertainty.

Chapters

Cited Sources

Concurring Sources

External References

Contribution & Novelties

The video provides a practitioner’s perspective on recent LLM architectures, highlighting practical considerations such as optimizer choices, expert collocation, and attention mechanisms. It offers a critical view of OpenAI’s open-weight model and discusses the trade-offs between width and depth.

Pour aller plus loin :

  • Muon optimizer — The optimizer used in Kimi K2, relevant to training stability.
  • DeepSeek-V3 — The architecture that Kimi K2 is based on, with details on expert collocation.
  • Attention Sinks — The concept of attention sinks, relevant to GPT-OSS’s bias approach.

85 words

Radar Profile

The radar profile shows high scores in technical level and information quality, reflecting the in-depth technical discussion. However, reliability is slightly lower due to the informal and speculative nature of the conversation.

Reliability 6/10