Le preguntamos a 15 IAs si matarían a 1000 personas

Le preguntamos a 15 IAs si matarían a 1000 personas

🎙 Codemancers - Inteligencia Artificial 👥 2K 📅 July 8, 2026 ⏱ 56 min 👁 190 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

Kobayashi MaruAI moral dilemmaLLM decision-makingAI safetybenchmark

Summary

In this episode of Codemancers, the hosts discuss the importance of university education in the AI era, arguing that knowledge is still valuable and that AI makes higher education more necessary, not less. They reference a Goldman Sachs executive’s claim that ‘knowledge is worth zero’ and the case of Finland’s PISA scores declining after shifting to reasoning-based teaching. The main segment presents the second edition of their ‘Kobayashi Maru’ moral benchmark, where they ask 15 AI models (including Opus 4.8, GPT 5.5, DeepSeek V4, and smaller models) to respond to a dilemma: whether to kill 1000 enemy soldiers to save millions. Results show that larger models now delegate the decision to a human, while smaller models (often used in edge devices) attack without hesitation, raising concerns about autonomous weapons. The hosts also discuss ‘Poison Fountain’, a Reddit community attempting to corrupt AI training data, and explain why it fails due to model collapse. They conclude that original creativity will gain value as AI becomes more prevalent.

166 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video offers a unique and timely perspective on AI moral reasoning, presenting an original benchmark that compares model behavior across sizes and versions. The hosts provide concrete examples and reasoning, making the argument engaging. However, the methodology is informal: the sample size is small, the prompt is not standardized across all models, and there is no statistical analysis. The discussion on university education is opinion-based, relying on anecdotal evidence and a single case study (Finland). The argument that AI makes university more necessary is plausible but not rigorously supported. The hosts do acknowledge the limitations of their benchmark, which adds credibility.

Scientific Rigor, Source Quality, Title Accuracy

The video cites several sources, but most are not formally referenced. The hosts mention a Goldman Sachs executive’s quote and Finland’s PISA results without providing specific citations. The benchmark results are presented as their own experiment, but no code or detailed methodology is shared (though they mention a GitHub repo ‘próximamente’). The description includes links to their podcast platforms and website, but no direct scientific sources. The title accurately reflects the content, and the video includes a clear structure with chapters. The hosts are transparent about the informal nature of their benchmark, which mitigates some concerns about rigor.

215 words

Title / Content Match

The title accurately reflects the core content: asking 15 AI models about a moral dilemma. The video delivers on this premise.

Quality & Reliability

6/10

The video presents an original moral benchmark on AI models, but the methodology is informal and lacks rigorous controls. The hosts provide subjective commentary and some unverified claims (e.g., Goldman Sachs executive quote, Finland PISA data). The benchmark results are anecdotal and not peer-reviewed.

Chapters

Cited Sources

Concurring Sources

  • AI Alignment — Supports the concern about AI moral decision-making and the need for alignment.

Dissenting Sources

  • No direct discordant sources found — The video's claims are not directly contradicted by any cited sources, but the lack of rigorous methodology means they should be treated with caution.

Contribution & Novelties

The video contributes an original, hands-on moral benchmark for AI models, comparing responses across generations and sizes. It highlights a concerning trend: smaller models, often deployed in edge devices, may make lethal decisions without human oversight. This is a valuable observation for AI safety discussions. The hosts also introduce the concept of ‘Poison Fountain’ and explain why such attacks are ineffective, adding to the discourse on data poisoning.

Pour aller plus loin :

118 words

Radar Profile

The radar profile shows moderate scores across all dimensions, with slightly higher quantity of information and lower technical depth. This reflects a balanced but informal presentation, suitable for a general audience interested in AI ethics.

Reliability 5/10

💬 No comments were provided for analysis.