J'étais SÛR et CERTAIN que CLAUDE 4 était surcoté et... WOW !

J'étais SÛR et CERTAIN que CLAUDE 4 était surcoté et... WOW !

🎙 Ludo Salenne 👥 267K 📅 May 26, 2025 ⏱ 42 min 👁 26K 📄 expert opinion 🧭 2026-08-21
Available in: English (current) Français

Keywords

Claude 4ChatGPTGeminicomparisonAI testing

Summary

In this video, Ludo Salenne conducts a practical comparison of three leading AI models: Claude 4 Sonnet, ChatGPT (likely GPT-4o), and Gemini 2.5 Pro. He performs four tests of increasing complexity: creating a landing page, developing a 3D RPG game, building an interactive dashboard, and generating an interactive report. For each test, he presents the results without revealing which AI produced them, allowing viewers to guess and then reveals the sources. The tests are designed to assess creativity, design, functionality, and adherence to instructions. Throughout the video, Ludo shares his subjective impressions, often favoring Claude 4 for its visual output and creativity. He also briefly tests Claude 4 Opus, noting its enhanced reasoning capabilities. The video concludes with Ludo’s overall verdict, where he expresses surprise at Claude 4’s performance, particularly for its free tier. The content is aimed at a general audience interested in AI applications, with a focus on practical use cases rather than technical depth.

157 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides a valuable, hands-on comparison of AI models in real-world scenarios, which is more relatable than abstract benchmarks. The argumentation is based on personal observation and subjective evaluation, which is clearly stated. The creator’s approach of blind testing adds an element of objectivity, but the final judgments are still influenced by personal preferences. The tests are well-chosen to cover different aspects of AI capabilities, from web design to game development. However, the lack of quantitative metrics and the small sample size limit the generalizability of the conclusions.

Scientific Rigor, Source Quality, Title Accuracy

The video does not cite external sources or studies; it relies solely on the creator’s own testing. The description includes links to the creator’s own resources and other videos, but no external references. The title is somewhat sensationalist but aligns with the content’s tone. The video’s scientific rigor is low, as it is an anecdotal comparison rather than a controlled study. The creator does not discuss potential biases or limitations of his methodology. The title accurately reflects the content but may overpromise a ‘WOW’ effect.

189 words

Title / Content Match

The title is catchy and reflects the creator's surprise at Claude 4's performance, but it is somewhat clickbait as it overstates the 'WOW' effect.

Quality & Reliability

6/10

The video is a subjective, hands-on comparison of three AI models based on four practical tests. The methodology is transparent but not rigorous (no control for variables, small sample, personal preference). The creator is transparent about his lack of prior experience with Claude 4. The claims are anecdotal and not backed by external benchmarks or peer-reviewed studies.

Chapters

Cited Sources

Concurring Sources

  • Claude 4 (Anthropic) — Official page for Claude 4, which the video claims to be superior in certain tasks.

Dissenting Sources

  • ChatGPT (OpenAI) — The video suggests ChatGPT underperforms in some tests, but OpenAI's official page claims high performance.

External References

Contribution & Novelties

The video offers a practical, user-centric comparison of AI models, which is more accessible than technical benchmarks. It highlights the importance of real-world testing over marketing claims. The ‘blind test’ format is an original approach to reduce bias.

Pour aller plus loin :

  • Claude 4 (Anthropic) — Official page for Claude 4, providing technical details and capabilities.
  • Gemini (Google) — Official page for Google’s Gemini AI.
  • ChatGPT (OpenAI) — Official page for ChatGPT.
  • AI Benchmarking — Wikipedia article on benchmarking, relevant to understanding how AI performance is measured.

88 words

Radar Profile

The radar profile shows moderate scores across all dimensions, indicating a balanced but not exceptional video. The highest score is in quantity of information, reflecting the multiple tests, while the lowest is in technical level, as the content is not deeply technical.

Reliability 5/10

💬 Très positif. Sur les 30 commentaires analysés, la grande majorité exprime des remerciements et des éloges pour la vidéo, soulignant son caractère pertinent et instructif. Quelques commentaires suggèrent des améliorations ou des tests supplémentaires, mais aucun ne remet en cause le contenu.