Je teste les IA ChatGPT o3-mini : Top ou Flop ?

Je teste les IA ChatGPT o3-mini : Top ou Flop ?

🎙 Ludo Salenne 👥 267K 📅 January 31, 2025 ⏱ 26 min 👁 43K 📄 expert opinion 🧭 2026-08-21
Available in: English (current) Français

Keywords

o3-minio3-mini-highChatGPTDeepSeekcomparatif

Summary

In this video, Ludo Salenne presents a first hands-on test of OpenAI’s new o3-mini and o3-mini-high models, released on January 31, 2025. He explains the differences between these models and the previous o1, notably the ability to connect to the internet but the lack of file attachment support. He conducts three blind tests comparing the two models on creative tasks: writing a comedic action script, debating a physics paradox, and pitching an absurd startup. He also tests the internet search feature and compares o3-mini with DeepSeek R1. The video highlights the message quotas for different subscription tiers, with o3-mini-high being unlimited only for Pro users at $200/month. The creator shares his subjective preferences and concludes that o3-mini-high generally performs better, but the launch seems rushed, possibly in response to Chinese AI competition. The video is informative for users wanting to understand the new models’ capabilities and limitations, but the evaluation is based on personal taste rather than rigorous benchmarks.

159 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides practical, hands-on information about the new o3-mini models, including access details, features, and limitations. The creator’s argumentation is based on personal testing and subjective preferences, which is transparent but not scientifically rigorous. He clearly states his opinion and encourages viewer feedback, making the evaluation interactive. The comparison with DeepSeek R1 is brief and not deeply analyzed, limiting the value of that section. Overall, the information is useful for users considering these models, but the lack of objective benchmarks weakens the argumentation.

Scientific Rigor, Source Quality, Title Accuracy

The video does not cite external scientific sources; the creator relies on his own experience and the model’s outputs. The description includes links to his own tutorials and promotional resources, which are not independent references. The title accurately reflects the content, as the video is a personal test and opinion. The creator acknowledges the rushed nature of the launch and the competitive pressure from Chinese AI, but does not provide evidence beyond his observations. The lack of rigorous methodology and external validation reduces the scientific reliability of the content.

188 words

Title / Content Match

The title accurately reflects the content: the creator tests ChatGPT o3-mini and o3-mini-high, giving a personal verdict. The title is slightly informal but matches the video's purpose.

Quality & Reliability

6/10

The video is a hands-on test and subjective comparison of AI models, based on personal experience and limited anecdotal evidence. The creator provides practical information about access and features, but the evaluation lacks rigorous methodology and relies on subjective preferences. The sources cited are mostly the creator's own tutorials and promotional links, not independent scientific references.

Chapters

Cited Sources

Concurring Sources

  • OpenAI o3-mini system card — Official documentation on o3-mini's capabilities and safety evaluations, consistent with the video's description of the model's features.

Dissenting Sources

  • DeepSeek-R1 technical report — The video briefly compares o3-mini with DeepSeek R1, but the technical report provides a more detailed and objective evaluation of DeepSeek's capabilities, which may differ from the creator's subjective impressions.

Contribution & Novelties

The video provides a timely, hands-on overview of the newly released o3-mini models, highlighting key features such as internet connectivity and the lack of file attachment support. It offers a practical comparison between o3-mini and o3-mini-high through creative tests, giving viewers a sense of their relative performance in informal tasks. The creator also discusses the competitive landscape with Chinese AI models like DeepSeek and Qwen, adding context to the release. However, the evaluation is subjective and lacks rigorous benchmarking, limiting its scientific contribution.

Pour aller plus loin :

  • OpenAI o3-mini system card — Official documentation on o3-mini’s capabilities and safety evaluations.
  • DeepSeek-R1 — Official page for DeepSeek-R1, a competing reasoning model.
  • Qwen 2.5 — Official blog post about Qwen 2.5, another Chinese AI model mentioned in the video.
  • Chain-of-thought prompting — Wikipedia article on chain-of-thought prompting, relevant to the reasoning capabilities of o3-mini.

143 words

Radar Profile

The radar profile shows moderate scores across all dimensions, with slightly higher scores in information quantity and technical level, reflecting the video's practical focus. The lower scores in information quality and reliability indicate the subjective and non-rigorous nature of the evaluation.

Reliability 5/10

💬 Très positif : Sur les 30 commentaires analysés, la majorité exprime une forte appréciation pour la réactivité et le contenu de la vidéo, avec de nombreux encouragements à prendre soin de sa santé. Quelques commentaires critiques soulignent que les tests ne sont pas représentatifs des capacités réelles des modèles, mais le ton général reste très favorable.