Gemini 3 vs GPT 5.1: La Comparativa Definitiva (Programación, Lógica y Visión)

Gemini 3 vs GPT 5.1: La Comparativa Definitiva (Programación, Lógica y Visión)

🎙 Codemancers - Inteligencia Artificial 👥 2K 📅 November 21, 2025 ⏱ 46 min 👁 264 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

Gemini 3GPT 5.1benchmarkprogrammingmultimodal

Summary

In this episode, the hosts compare Google’s Gemini 3 with OpenAI’s GPT 5.1 through a series of practical tests. They start with logic and vision tasks, such as identifying the largest candle from ASCII art, reading ancient Catalan text, counting arms in an image, and telling time from a watch. Gemini 3 often provides more detailed explanations and correctly identifies the watch model, while GPT 5.1, configured for conciseness, gives shorter answers. In web development, both generate functional pages, but Gemini 3 excels in creating a data dashboard from an Excel file on the first try, whereas GPT 5.1 requires iteration. The definitive test involves programming a maze game in a single HTML file; Gemini 3 produces a working game with learning capabilities, while GPT 5.1 fails to generate a functional result. The hosts conclude that Gemini 3 is technically superior, especially in multimodal and programming tasks, but note that for average users, the differences are minimal. They also discuss the flattening innovation curve and the niche advantages in programming, where Claude remains strong. The episode includes a segment on GitHub Copilot’s resurgence, its cost-effectiveness, and its role as a backup tool alongside Claude Code.

195 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable hands-on comparisons of two leading AI models, offering practical insights into their real-world performance. The argumentation is based on direct testing, which adds credibility, but it is limited by the informal methodology and the hosts’ personal configurations. The hosts acknowledge potential biases, such as GPT 5.1 being ’tuned’ for conciseness, which affects the comparison. The discussion on the flattening innovation curve and the niche advantages in programming is insightful, though it relies on subjective impressions rather than systematic benchmarks. Overall, the value lies in the practical demonstrations and the honest assessment of the models’ strengths and weaknesses.

Scientific Rigor, Source Quality, Title Accuracy

The video does not cite external scientific sources; it relies on the hosts’ own testing and observations. The title accurately reflects the content, which is a comparative analysis. The hosts mention benchmarks but do not provide specific data or references. The lack of rigorous methodology and the absence of verifiable sources reduce the scientific rigor. However, the practical tests are transparent and reproducible, which partially compensates for the lack of formal citations. The title is appropriate and not misleading.

195 words

Title / Content Match

The title accurately reflects the content, which is a detailed comparison of Gemini 3 and GPT 5.1 across programming, logic, and vision tasks.

Quality & Reliability

6/10

The video presents a hands-on comparison of two AI models, with practical tests and personal observations. However, the methodology is informal, lacks rigorous controls, and relies on anecdotal evidence. The hosts acknowledge their own biases and configuration differences, which limits the reliability of the conclusions.

Chapters

Cited Sources

Contribution & Novelties

The video offers a practical, side-by-side comparison of Gemini 3 and GPT 5.1, focusing on programming, logic, and vision tasks. It provides concrete examples of their performance, highlighting Gemini 3’s strengths in multimodal understanding and code generation. The discussion on the flattening innovation curve and the niche advantages in programming adds a valuable perspective. However, the lack of rigorous methodology and the reliance on anecdotal evidence limit the novelty and generalizability of the findings.

Pour aller plus loin :

  • Gemini 3 official page — Official information about Gemini 3.
  • GPT-5.1 official page — Official information about GPT-5.1.
  • SWE-bench benchmark — A benchmark for evaluating AI models on real-world software engineering tasks, relevant to the programming comparison.

116 words

Radar Profile

The radar profile shows moderate scores across all dimensions, with quantity of information and technical level slightly higher than quality and reliability. This suggests the video provides a decent amount of technical detail but lacks rigorous sourcing and methodological depth.

Reliability 5/10