
Probamos GPT-5.2 vs Gemini en escenarios reales. Estos son los resultados
Keywords
Summary
166 words
Critical Evaluation
The video provides a reasonably informative comparison of GPT-5.2 and Gemini 3, but its scientific rigor is limited. The presenter relies heavily on official benchmarks from OpenAI, which may be biased, and acknowledges this by noting that companies often fine-tune models for specific tests. However, he does not critically evaluate the benchmarks’ validity or independence. The hands-on tests are subjective and not controlled, with only a few examples, making it difficult to draw general conclusions. The presenter’s expertise is not formally established, and the video includes a promotional segment for EDteam, which could introduce bias. The analysis of the ocean wave simulation is illustrative but not systematic. The video does not cite external sources or peer-reviewed studies, and the sources listed in the description are mostly promotional links. Overall, the video offers a useful overview for a general audience but lacks the depth and rigor expected in a scientific evaluation. The adéquation titre/contenu is good, as the title accurately reflects the content. The video’s strength lies in its practical demonstrations, but its weakness is the lack of independent verification and the potential for bias due to the presenter’s affiliation with EDteam and the promotional content.
195 words
Title / Content Match
The title accurately reflects the content, which is a comparative test of GPT-5.2 and Gemini 3 in real-world scenarios.
Quality & Reliability
6/10
The video provides a comparative analysis of GPT-5.2 and Gemini 3, based on official benchmarks and hands-on tests. However, the analysis is largely subjective, lacks independent verification, and relies on promotional material from OpenAI. The presenter's expertise is not formally established, and the video includes a promotional segment.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction: GPT-5.2 as a response to Gemini 3
- Context: Google's rise and OpenAI's code red
- Overview of GPT-5.2 features and benchmarks
- Explanation of GDP Val benchmark and human expert comparison
- Example tasks: workforce planning and cap table analysis
- Comparison of GPT-5.1 vs GPT-5.2 on complex tasks
- SWE-Bench Pro results and programming comparison
- Hands-on test: ocean wave simulation
- Hands-on test: party card creator and typing game
- Comparison of results and conclusion
Cited Sources
- EDteam — Platform for programming and AI courses, mentioned in the video.
- EDteam free courses — Link to free courses, mentioned in the video.
- EDteam courses — Link to all courses, mentioned in the video.
- EDteam scholarships — Link to scholarships, mentioned in the video.
- EDteam Instagram — Social media link, mentioned in the video.
- EDteam LinkedIn — Social media link, mentioned in the video.
- EDteam premium — Link to premium subscription, mentioned in the video.
- EDteam teachers — Link for teachers, mentioned in the video.
- EDteam TikTok — Social media link, mentioned in the video.
Concurring Sources
- OpenAI GPT-5.2 announcement — Official announcement of GPT-5.2, supporting the video's claims about its capabilities.
- Google Gemini 3 announcement — Official announcement of Gemini 3, providing context for the comparison.
Dissenting Sources
- Independent benchmark evaluations — The video relies on OpenAI's own benchmarks, which may be biased. Independent evaluations from third parties are not mentioned.
Contribution & Novelties
The video provides a practical, hands-on comparison of GPT-5.2 and Gemini 3, going beyond official benchmarks. It highlights GPT-5.2’s strengths in productivity tasks like spreadsheet and presentation creation, while acknowledging Gemini 3’s superiority in image generation. The presenter’s real-world tests, such as the ocean wave simulation, offer tangible insights into the models’ capabilities. However, the analysis is not exhaustive and lacks independent verification.
Pour aller plus loin :
- GPT-5.2 official page — Official information about GPT-5.2.
- Gemini 3 official page — Official information about Gemini 3.
- SWE-bench — Benchmark for software engineering tasks, relevant to the programming comparison.
- GDPval — OpenAI’s benchmark for professional tasks, mentioned in the video.
109 words
Radar Profile
The radar profile shows moderate scores across all dimensions, with quantity of information being the highest (7) and technical level the lowest (5). This suggests the video provides a good amount of information but lacks deep technical analysis, making it suitable for a general audience rather than experts.
💬 No comments were provided for analysis.