
BEST Local AI for Tool Calls? GPT-OSS, DeepSeek, Kimi K2, MiniMax, GLM Compared
Keywords
Summary
138 words
Critical Evaluation
Value of the Information & Strength of the Argument
The video provides valuable hands-on insights into the tool calling capabilities of local AI models, which are often underrepresented in benchmarks. The empirical tests, though informal, reveal real-world failure modes like rate limiting and hallucinated URLs, and demonstrate how models adapt. The argumentation is coherent, with clear comparisons and visual evidence. However, the lack of standardized metrics, single-run tests, and potential biases in hardware/software configuration limit the strength of conclusions. The host’s reasoning is transparent, and he acknowledges limitations, but the subjective evaluation of ‘smartness’ is not reproducible.
Scientific Rigor, Source Quality, Title Accuracy
The video references model cards on Hugging Face and the Inferencer app, providing traceability. The testing procedure is not controlled, but the host shows actual outputs and errors, enhancing credibility. The title is apt, as it accurately reflects the content. However, there is no external validation, and the affiliate links in the description introduce potential conflicts of interest. The absence of citations to peer-reviewed benchmarks reduces scientific rigor, but as an expert opinion piece, it offers practical value.
181 words
Title / Content Match
The title accurately describes the content, as the video systematically compares several local AI models on their tool calling capabilities, directly matching the viewer's expectation.
Quality & Reliability
6/10
The video is an informal, hands-on comparative test of local AI models for tool calling. While the demonstrations are real and reproducible, the methodology is not rigorously controlled, lacks statistical validation, and includes potential biases from affiliate marketing. However, the host transparently shows limitations and model behaviors, grounding the assessment in concrete outputs.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction to local AI tool calling and the Inferencer app.
- Demonstration of the weather tool call with a live request.
- Testing search limitations with GPT-OSS 20B and 120B.
- Comparing MiniMax M2 and Kimi K2 handling of truncated content.
- DeepSeek's prompt-based tool calling and adaptation.
- Creating a custom tool using Qwen3 to modify search provider.
- Testing Apertus and final summary of model performance.
Cited Sources
- GPT-OSS 20B model — One of the tested models for tool calling.
- GPT-OSS 120B model — One of the tested models for tool calling.
- DeepSeek V3.1 Terminus MLX — Tested model with prompt-based tool calling.
- GLM 4.6 MLX — Tested model for tool calling.
- Kimi K2 Thinking MLX — Tested model known for strong reasoning.
- MiniMax M2 MLX — Tested model with XML-based tool call format.
- Qwen3 1.7B — Small model tested for tool calling on resource-constrained devices.
- Qwen3 235B A22B Instruct — Large Qwen model tested for tool calling.
- Apertus 70B — Swiss model tested, showing basic tool calling.
- Inferencer App — Application used to run local models and manage tool calls.
External References
Contribution & Novelties
The video offers a practical comparison of tool calling across multiple open-weight local AI models, revealing that while all can call basic functions, their robustness varies significantly under constraints like rate-limited searches. It demonstrates GPT-OSS 120B’s ability to creatively bypass failed search tools, and highlights the XML parsing issues in MiniMax. The host also shows how to easily create custom tools within Inferencer, enhancing extendibility. This hands-on approach provides insights that benchmarks often miss, particularly regarding real-world failure recovery.
Pour aller plus loin :
- Toolformer: Language Models Can Teach Themselves to Use Tools — A foundational paper on tool usage in LLMs, providing context for the tool calling mechanisms tested.
- ReAct: Synergizing Reasoning and Acting in Language Models — Explains the reasoning-and-acting loop that underlies agentic tool use, similar to the behavior demonstrated.
- Function Calling in LLMs (OpenAI) — A widely adopted standard for tool definitions and outputs, mentioned as a comparison in the video.
155 words
Radar Profile
The radar indicates a high quantity and quality of information, with a moderate level of technical depth and a lower global reliability due to the informal methodology. The video excels in providing practical demonstrations and diverse model comparisons, but the lack of controlled experiments and statistical rigor pulls down the reliability score.