Let's Run Qwen 3.5 - Local AI HERO Model for OpenClaw, Writing, Coding & More

Let's Run Qwen 3.5 - Local AI HERO Model for OpenClaw, Writing, Coding & More

🎙 xCreate 👥 26K 📅 February 16, 2026 ⏱ 26 min 👁 19K 📄 tutorial 🧭 2026-09-09
Available in: English (current) Français

Keywords

Qwen 3.5local inferencetoken speedagent toolsbenchmark

Summary

This video presents a comprehensive hands-on evaluation of the newly released Qwen 3.5 open-weight model (397B parameters, 17B active) running locally on an M3 Ultra Mac Studio with 512GB RAM. The host tests output speed (~24 tokens/s), demonstrates batching with multiple concurrent inferences, verifies tool calling and grounding via Wikipedia, integrates with OpenClaw for agent tasks (web search, Apple Notes), and uses Kilo Code for coding tasks. Creative writing and logical reasoning are also assessed, including tests with thinking disabled/enabled. Practical configurations like prompt caching and quantization (Q9) are explained, along with notes on memory requirements (~415GB) and potential US export bans. The video concludes with a positive appraisal of the model’s creative and general capabilities, while noting minor issues in coding tasks like copying existing code.

127 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides concrete, reproducible data points: token speeds, memory usage, and completion quality across various tasks. The presenter grounds his claims through actual runs, comparing Qwen 3.5 against models like Kimi K2.5 and GLM, and highlights both strengths (creative writing, agent integration) and weaknesses (coding plagiarism, logic puzzles with thinking disabled). The argumentation is logical and transparent about hardware requirements, quantization trade-offs, and the need for enabling thinking for reliable tool calls. However, the absence of independent benchmark verification and the host’s admitted bias toward Qwen slightly weaken the objectivity.

Scientific Rigor, Source Quality, Title Accuracy

The video cites Hugging Face and Inferencer as primary references, and provides companion video links to other model evaluations. The presenter does not cite academic papers, relying instead on empirical testing, which is acceptable for a practical tutorial. The title accurately reflects the content, promising a local run of Qwen 3.5, which is fulfilled. The description includes affiliate links, but they do not affect the technical content. The host is transparent about assumptions and limitations, such as the use of Q9 quants and emulation for FP8. Overall, the information is presented with reasonable rigor, but the lack of external verification of claimed benchmarks restricts the score.

212 words

Title / Content Match

The title accurately reflects the content: running Qwen 3.5 locally, with emphasis on its performance for agent tasks, creative writing, and coding.

Quality & Reliability

7/10

The video provides a hands-on demonstration with multiple benchmarks, real-time token rates, and integration tests (OpenClaw, Kilo Code). The presenter is transparent about limitations and quantization choices, but the content includes affiliate links and the benchmark sources are not independently verified.

Key Moments

Cited Sources

Concurring Sources

External References

Contribution & Novelties

This video offers an original, real-time assessment of Qwen 3.5’s local performance, particularly its 397B-parameter open-weight architecture with only 17B active parameters, making it feasible for high-end consumer hardware. It highlights the model’s creative writing capabilities and its integration with agentic tools like OpenClaw, while also exposing practical quirks like prompt caching and quantization trade-offs. The host’s candid discussion of failures (e.g., coding plagiarism) provides useful insights for the community.

Pour aller plus loin :

  • Qwen official — The official model family page with documentation and variations.
  • Mixture of Experts — Background on the MoE architecture that Qwen 3.5 uses.
  • Prompt Caching — Explanation of how caching reduces latency in LLM inference.
  • Quantization in Large Language Models — Notes on quantizing models for local deployment (no URL if uncertain, but this is a known blog).

135 words

Radar Profile

The radar profile is balanced, with high scores in information quantity and practicality, and moderate scores in quality/rigor. The low level of technical depth compared to research papers is offset by the step-by-step tutorial nature.

Reliability 7/10