J'ai testé GPT-5 : vous aviez raison d'être sceptiques ...

J'ai testé GPT-5 : vous aviez raison d'être sceptiques ...

🎙 Vision IA 👥 294K 📅 August 12, 2025 ⏱ 27 min 👁 40K 📄 expert opinion 🧭 2026-08-21
Available in: English (current) Français

Keywords

GPT-5AI comparisoncodingbenchmarkOpenAI

Summary

The video presents a comprehensive, hands-on comparison of GPT-5 against seven other AI models (Gemini 2.5 Pro, Claude, Grok, Kimi K2, DeepSeek, Mistral, and Qwen) across eight technical challenges: a logic puzzle, a Mario game, a Minecraft clone, a spreadsheet application, a music synthesizer, a shader editor, a 3D racing game, and a meal planner. The creator, Vision IA, aims to provide factual evidence to address the polarized reception of GPT-5. Each test is conducted with the same prompt, and results are shown live. The findings reveal that GPT-5 excels in some creative coding tasks (e.g., the 3D racing game) but fails in others (e.g., Minecraft, spreadsheet), confirming its instability. The video also highlights the strengths of competitors: Gemini 2.5 Pro performs well in Minecraft and spreadsheet, Claude impresses with the shader editor and meal planner, and Kimi K2 shines in the synthesizer test. The creator concludes that GPT-5’s performance is inconsistent, likely due to its routing system, and advises users to choose models based on specific needs. The video ends with a call to subscribe and join the community.

180 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable, practical insights into the real-world performance of GPT-5 and its competitors. The argumentation is solid: the creator uses a structured, multi-test methodology, showing prompts and results transparently. The tests cover a range of difficulty levels and domains, from logic to complex coding, which strengthens the validity of the conclusions. However, the evaluation is based on a single tester’s experience, which may introduce bias. The creator acknowledges this limitation and encourages viewers to test themselves. The argumentation is convincing in demonstrating GPT-5’s instability, but it could be strengthened by including more quantitative metrics (e.g., success rates, time to completion) and by testing a larger sample of prompts.

Scientific Rigor, Source Quality, Title Accuracy

The video demonstrates a reasonable level of scientific rigor: the methodology is clear, and the results are presented without excessive hype. The creator does not cite external sources, but the video is based on direct experimentation, which is appropriate for a hands-on test. The title accurately reflects the content, and the video’s structure (with timestamps) aids navigation. However, the lack of peer review and the subjective nature of the tests limit the overall reliability. The creator’s expertise in AI is evident, but the video would benefit from including more context on the models’ versions and settings. The description provides links to the creator’s newsletter and training, but these are promotional and not scientific sources.

239 words

Title / Content Match

The title accurately reflects the content: a skeptical, hands-on test of GPT-5's capabilities, confirming user concerns about its instability.

Quality & Reliability

7/10

The video provides a structured, hands-on comparison of GPT-5 against seven other AI models across eight technical challenges. The methodology is transparent (prompts shown, results displayed), but it is based on a single tester's experience and lacks statistical rigor or peer review. The creator's expertise in AI is evident, but the assessment is subjective and may not generalize.

Key Moments

Cited Sources

Concurring Sources

  • OpenAI GPT-5 official page — Official information about GPT-5's capabilities and limitations.

Dissenting Sources

  • User reports of GPT-5 being powerful — Some users report GPT-5 being powerful, but the video shows inconsistencies.

Contribution & Novelties

The video offers a unique, hands-on comparison of GPT-5 against seven other AI models across eight diverse coding challenges, providing concrete evidence of its strengths and weaknesses. Unlike typical benchmark reports, this test focuses on practical, user-relevant tasks, making the findings accessible and actionable. The video also highlights the importance of considering model routing and version differences, which are often overlooked in public discussions.

Pour aller plus loin :

  • OpenAI GPT-5 official page — Official information about GPT-5’s capabilities and limitations.
  • Benchmarking LLMs: A Comprehensive Guide — Academic paper on LLM evaluation methodologies.
  • Model Routing in AI Systems — Overview of routing techniques used in AI systems, relevant to GPT-5’s inconsistent performance.

112 words

Radar Profile

The radar profile shows high scores in information quantity and technical level, reflecting the video's comprehensive and detailed testing. However, the lower scores in information quality and reliability indicate that the subjective nature of the tests and lack of external validation limit the overall robustness. The video is informative but should be complemented with more rigorous, peer-reviewed evaluations.

Reliability 6/10

💬 Équilibré. Sur les 30 commentaires analysés, les avis sont partagés : certains saluent la qualité du test et la transparence, tandis que d'autres expriment leur déception vis-à-vis de GPT-5 et demandent plus de tests sur des usages courants. Plusieurs commentaires signalent un problème de son à la fin de la vidéo.