J'ai refait Pokémon avec Claude Opus 4.8 en un seul prompt

J'ai refait Pokémon avec Claude Opus 4.8 en un seul prompt

I remade Pokémon with Claude Opus 4.8 in a single prompt

🎙 iAlan 👥 8K 📅 May 30, 2026 ⏱ 15 min 👁 1K 📄 expert opinion 🧭 2026-09-05
Available in: English (current) Français

Keywords

Claude Opus 4.8UltraCodeGPT-5.5Qwen 3.7 MaxPokémon

Summary

The video presents a practical test of Claude Opus 4.8, a new AI model from Anthropic, focusing on its coding capabilities. The creator compares it with GPT-5.5 and Qwen 3.7 Max on several tasks: a 2D game, a Windows clone, a perfume landing page, and a 3D flight simulator. The results show that Opus 4.8 performs well, but not always the best; Qwen 3.7 Max surprisingly excels in the 2D game and flight simulator. The creator then demonstrates the new UltraCode mode in Claude Code, which spawns hundreds of sub-agents to autonomously plan and execute tasks. He uses this to recreate a Pokémon game in a single prompt, achieving an impressive result with a playable intro and character selection, but the game is incomplete and the cost is high (15€). The video also mentions the upcoming Claude Mythos model and provides practical tips for using Claude Code. Overall, the video is informative for AI enthusiasts interested in hands-on testing, but it lacks rigorous scientific methodology.

165 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable hands-on insights into the practical capabilities of Claude Opus 4.8, especially its new UltraCode mode. The creator demonstrates real-world performance on diverse tasks, which is more informative than benchmark numbers alone. The argumentation is based on direct observation and comparison, but it is subjective and lacks controlled conditions. The creator acknowledges limitations, such as the high token consumption and incomplete results, which adds credibility. However, the evaluation criteria are not clearly defined, and the conclusions are based on personal preference rather than objective metrics.

Scientific Rigor, Source Quality, Title Accuracy

The video cites official claims from Anthropic about reduced hallucinations and improved efficiency, but does not provide direct links to these sources. The creator mentions benchmarks like SWE-bench and Terminal Bench, but does not give specific URLs. The description includes a link to a form for prompts, which is not a scientific source. The title accurately reflects the content, focusing on the Pokémon recreation, which is the highlight. The video is more of a practical demonstration than a rigorous scientific analysis, so the scientific rigor is moderate.

190 words

Title / Content Match

The title accurately reflects the main content: the creator uses Claude Opus 4.8 in a single prompt to recreate a Pokémon game, with additional comparative tests.

Quality & Reliability

6/10

The video is a hands-on test of Claude Opus 4.8, providing practical demonstrations and comparisons with other models. However, the methodology is informal, lacks rigorous controls, and relies on subjective impressions. The creator mentions benchmarks and official claims but does not provide detailed sources or verification.

Chapters

Cited Sources

  • Prompts pack (free download) — The creator offers the prompts used in the tests for free via this form.

Concurring Sources

  • Anthropic's official claims about Opus 4.8 — The creator mentions Anthropic's claims about reduced hallucinations and improved efficiency, which are consistent with the video's observations.

Dissenting Sources

  • Benchmark comparisons — The video shows that Opus 4.8 does not always outperform GPT-5.5 or Qwen 3.7 Max, which contrasts with some benchmark rankings.

Contribution & Novelties

The video offers a practical, real-world evaluation of Claude Opus 4.8, particularly its new UltraCode mode, which is a significant feature. The creator demonstrates that the model can autonomously plan and execute complex tasks, such as recreating a Pokémon game, with impressive results. This provides valuable insights for developers and AI enthusiasts. The comparison with other models adds perspective, though the methodology is informal.

Pour aller plus loin :

108 words

Radar Profile

The radar profile shows a balanced but moderate performance across all dimensions, with slightly higher scores in information quantity and technical level, reflecting the hands-on nature of the video. The lower scores in information quality and reliability indicate the lack of rigorous methodology and sources.

Reliability 5/10

💬 Sur les 0 commentaires analysés, aucune tendance n'est disponible.