I was WRONG about M5 MacBook Pro 🤯 (vs M4 Max & M3 Ultra) for Local AI

I was WRONG about M5 MacBook Pro 🤯 (vs M4 Max & M3 Ultra) for Local AI

🎙 xCreate 👥 26K 📅 April 14, 2026 ⏱ 13 min 👁 21K 📄 expert opinion 🧭 2026-09-09
Available in: English (current) Français

Keywords

prompt processingtokens per secondbatchingdistributed inferenceneural accelerators

Summary

The video is a comparative review of the 2026 M5 Pro MacBook Pro against the M4 Max and M3 Ultra, focusing on local AI/LLM performance. The creator tests Qwen 3.5 27B using a local inference app, measuring prompt processing time, token generation speed, batching performance, and distributed compute across machines. Key results show that the M5 Pro’s prompt processing is about twice as fast as the M4 Max (8 s vs 17 s) and even beats the M3 Ultra (9.37 s), thanks to the new NAX neural accelerators. Generation speed, however, remains higher on the M4 Max (25 vs 17.5 tokens/s). Batching tests show the M4 Max reaches higher aggregated throughput (up to 52 tokens/s) than the M5 Pro (32 tiles/s), though the M5 Pro’s per-stream speed is consistent. The author also demonstrates successful distributed inference between the M5 Pro and M4 Max, confirmed with a remote server setup. The video concludes that the M5 Pro is an excellent purchase, especially for prompt-heavy workloads, and anticipates a future M5 Max with even greater gains. The creator openly admits he was wrong about the magnitude of the improvement, and praises Apple’s implementation of the neural accelerators.

195 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable, concrete benchmark data on a newly released chip, filling a gap in independent performance assessments for local AI use. The argumentation is strengthened by the author’s willingness to revisit his own earlier predictions and by the transparency regarding potential confounds (e.g., prompt caching). The testing procedure is clearly described, with the model, quantisation, and prompt size specified, and the use of OBS screen recording is acknowledged as a possible influence. The distinction between prompt processing and generation speed is explained clearly, and the author discusses the practical implications for different types of AI workloads. The batching and distributed-computing tests add depth, showing real-world scenarios. However, the argumentation relies solely on the author’s own measurements without cross-validation from other sources, and no statistical analysis is performed, limiting the quantitative reliability.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigour is moderate: the methodology is reproducible in principle, but no official documentation from Apple or independent academic sources is cited; all data are self-generated. The sources listed in the description are primarily affiliate links or companion videos, none of which provide external validation. The title is accurate and does not exaggerate the findings, and the content matches the description. The author’s candidness about the prompt-caching ‘cheat’ enhances credibility. No attended public reactions are analysed here, but the comments (provided separately) suggest a generally favourable reception, with some requests for clearer visual representation of results.

245 words

Title / Content Match

The title accurately reflects the content, as the author openly retracts a previous assertion about the M5's prompt processing improvements.

Quality & Reliability

8/10

The video is methodologically sound, with transparent benchmarks, voluntary disclosure of a confounding variable (prompt caching), and appropriate acknowledgment of uncertainty in the author's initial predictions. However, it represents a single individual's testing rather than a peer-reviewed study, so the score is slightly reduced.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The video delivers one of the first independent, hands-on benchmarks for the M5 Pro’s NAX neural accelerators in the context of LLM inference. It quantifies the dramatic prompt-processing speedup (about 4x over previous generation) and shows that even the Pro variant outperforms the M3 Ultra in this metric. The author’s admission of error adds originality by demonstrating that initial conjectures can be overturned with empirical data. Practical insights for developers and buyers are provided, such as the importance of prompt processing in agentic workloads and the feasibility of distributed inference.

Pour aller plus loin :

  • Apple Neural Engine — Overview of Apple’s dedicated hardware for machine learning, relevant to understanding the NAX accelerators.
  • LLM inference optimization — Research on reducing latency and improving throughput in language model inference, complementing the video’s findings.
  • Distributed inference — General concept of distributing computation across nodes, applied here to run a single LLM across two Macs.

152 words

Radar Profile

The radar profile shows high scores in information quality and technical depth, slightly lower in quantity and reliability. This reflects a focused, well-executed test but limited scope and lack of external validation, typical of channel-based reviews.

Reliability 7/10

💬 Orientation: positif. Sur les 30 commentaires analysés, ils sont globalement favorables, mais plusieurs demandent des tableaux de chiffres plus lisibles et soulignent une confusion entre les noms de modèles et de puces.