
Le preguntamos a 15 IAs si matarían a 1000 personas
Keywords
Summary
166 words
Critical Evaluation
Value of the Information & Strength of the Argument
The video offers a unique and timely perspective on AI moral reasoning, presenting an original benchmark that compares model behavior across sizes and versions. The hosts provide concrete examples and reasoning, making the argument engaging. However, the methodology is informal: the sample size is small, the prompt is not standardized across all models, and there is no statistical analysis. The discussion on university education is opinion-based, relying on anecdotal evidence and a single case study (Finland). The argument that AI makes university more necessary is plausible but not rigorously supported. The hosts do acknowledge the limitations of their benchmark, which adds credibility.
Scientific Rigor, Source Quality, Title Accuracy
The video cites several sources, but most are not formally referenced. The hosts mention a Goldman Sachs executive’s quote and Finland’s PISA results without providing specific citations. The benchmark results are presented as their own experiment, but no code or detailed methodology is shared (though they mention a GitHub repo ‘próximamente’). The description includes links to their podcast platforms and website, but no direct scientific sources. The title accurately reflects the content, and the video includes a clear structure with chapters. The hosts are transparent about the informal nature of their benchmark, which mitigates some concerns about rigor.
215 words
Title / Content Match
The title accurately reflects the core content: asking 15 AI models about a moral dilemma. The video delivers on this premise.
Quality & Reliability
6/10
The video presents an original moral benchmark on AI models, but the methodology is informal and lacks rigorous controls. The hosts provide subjective commentary and some unverified claims (e.g., Goldman Sachs executive quote, Finland PISA data). The benchmark results are anecdotal and not peer-reviewed.
Chapters
- Intro: benchmark moral + una comunidad rarísima de Reddit
- Fable se va el 12 (y es el único modelo sin zero-data-retention)
- Truco: usa Fable como orquestador que escribe el roadmap
- El "full stack chino": planificar caro, ejecutar barato
- ¿Sirve ir a la universidad en la era de la IA?
- "El conocimiento vale cero" (el ejecutivo de Goldman Sachs)
- El caso de Finlandia y los exámenes PISA
- Por qué la IA hace MÁS necesaria la universidad
- Coste de oportunidad: elige una carrera con salidas
- El benchmark Kobayashi Maru: qué es
- El prompt exacto del dilema del dron
- Resultados de hace un año (2024)
- Resultados hoy (2026): los grandes delegan en el humano
- Los modelos pequeños atacan sin pensar (el peligro del edge)
- Amália y Rio 3.5: los modelos de gobiernos
- Sesgo étnico: el experimento Israel/Palestina
- Cómo puntuar el benchmark (y el repo en GitHub)
- Poison Fountain: la secta que quiere envenenar a la IA
- Por qué no funciona: perplejidad y colapso del modelo
- La creatividad original ganará valor
- Cierre (y maldito Streamlabs)
Cited Sources
- Codemancers Podcast on Spotify — Podcast platform where the episode is available.
- Codemancers Podcast on Apple Podcasts — Podcast platform where the episode is available.
- Codemancers Website — Official website of the podcast.
Concurring Sources
- AI Alignment — Supports the concern about AI moral decision-making and the need for alignment.
Dissenting Sources
- No direct discordant sources found — The video's claims are not directly contradicted by any cited sources, but the lack of rigorous methodology means they should be treated with caution.
Contribution & Novelties
The video contributes an original, hands-on moral benchmark for AI models, comparing responses across generations and sizes. It highlights a concerning trend: smaller models, often deployed in edge devices, may make lethal decisions without human oversight. This is a valuable observation for AI safety discussions. The hosts also introduce the concept of ‘Poison Fountain’ and explain why such attacks are ineffective, adding to the discourse on data poisoning.
Pour aller plus loin :
- Kobayashi Maru (Star Trek) — The original fictional test, providing context for the benchmark.
- AI alignment — The field concerned with ensuring AI systems act in accordance with human values.
- Lethal autonomous weapons — Discussion on the ethical and safety implications of autonomous weapons systems.
118 words
Radar Profile
The radar profile shows moderate scores across all dimensions, with slightly higher quantity of information and lower technical depth. This reflects a balanced but informal presentation, suitable for a general audience interested in AI ethics.
💬 No comments were provided for analysis.