
Why everyone HATES GPT-5 (and how to fix it)
Keywords
Summary
179 words
Critical Evaluation
Value of the Information & Strength of the Argument
The video provides substantial value by directly testing the most common complaints against GPT-5, offering concrete examples and comparisons. Wolfe’s approach is methodical: he recreates user-reported issues, compares outputs across multiple models, and evaluates the validity of each gripe. His argumentation is balanced, acknowledging both valid criticisms (e.g., coding performance, accuracy) and those he considers less valid (e.g., personality, based on his tests). He also contextualizes the model routing issue, explaining how a broken router could have made GPT-5 seem dumber. However, the analysis is limited by its anecdotal nature; the tests are not comprehensive or statistically rigorous, and personal experience may not generalize. Wolfe also speculates on OpenAI’s motivations (e.g., cost-saving) without direct evidence, which weakens the argumentative rigor.
Scientific Rigor, Source Quality, Title Accuracy
The video references several sources, including Reddit posts, tweets from Sam Altman, and comments from AI experts like Ethan Mollick, but these are presented without formal citations or links in the description. The description includes links to FutureTools.io, a newsletter, and a sponsor (Globant), but no direct references to the cited claims. The title accurately reflects the content, which is a critical review of GPT-5. The video’s scientific rigor is moderate: it uses real-world tests but lacks controlled conditions and peer-reviewed sources. The adéquation between title and content is strong, as the video indeed explores why people hate GPT-5 and suggests fixes. However, the lack of verifiable sources for the claims made (e.g., benchmark charts) reduces the overall reliability.
254 words
Title / Content Match
The title accurately reflects the content, which systematically addresses common criticisms of GPT-5 and explores potential fixes.
Quality & Reliability
7/10
The video provides a balanced, hands-on evaluation of user complaints about GPT-5, testing claims with direct experiments and referencing community feedback. However, the analysis is anecdotal and lacks rigorous methodology or independent verification.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
Cited Sources
- FutureTools.io — Matt Wolfe's AI tools and news platform, mentioned as a resource for exploring AI tools.
- Weekly Newsletter — Sign-up for Matt Wolfe's weekly AI newsletter.
- Globant Enterprise AI — Sponsor segment promoting Globant's AI services, not directly related to the video's content.
- Matt Wolfe's LinkedIn — Social media profile for Matt Wolfe.
- Matt Wolfe's Threads — Social media profile for Matt Wolfe.
Concurring Sources
- Reddit post on GPT-5 accuracy — User-reported logic puzzle failure, consistent with the video's test results.
- Sam Altman's tweet on router issue — Acknowledgment that the auto-switcher broke, supporting the video's explanation for perceived dumbing down.
Dissenting Sources
- OpenAI's official benchmark claims — OpenAI's published benchmarks suggest GPT-5 is superior, but the video's tests show mixed results, indicating a discrepancy between official claims and real-world performance.
Contribution & Novelties
The video offers a practical, hands-on evaluation of GPT-5’s shortcomings, going beyond surface-level complaints by testing specific claims. It provides a useful framework for users to assess AI model updates and offers actionable advice, such as enabling legacy models in settings. The comparison across multiple models (GPT-4o, GPT-3 Pro, Claude Opus 4.1, Grok 4) adds value for users deciding which AI to use.
Pour aller plus loin :
- OpenAI’s official blog on GPT-5 — Official information about GPT-5’s capabilities and limitations.
- Ethan Mollick’s Substack — Analysis of AI models and their practical implications, referenced in the video.
- Balatro game — The game used in the coding test, for context.
- AI model benchmarks on Papers with Code — A resource for comparing AI model performance on standardized benchmarks.
127 words
Radar Profile
The radar profile shows high scores in information quantity and quality, reflecting the video's comprehensive coverage and practical tests. However, technical depth and reliability are moderate, as the analysis is anecdotal and lacks rigorous methodology. The overall balance suggests a useful but not deeply scientific review.
💬 Négatif : Sur les 30 commentaires analysés, la grande majorité exprime une forte insatisfaction envers GPT-5, notamment sur la perte de personnalité, la baisse de qualité et les problèmes de fiabilité, avec quelques voix qui apprécient les solutions proposées.