MIT 6.S191: The Three Laws of AI

MIT 6.S191: The Three Laws of AI

🎙 Doug Blank 👥 356K 📅 May 11, 2026 ⏱ 51 min 👁 10K 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

LLMevaluationjailbreakpromptsafety

Summary

Doug Blank, head of research at Comet ML, presents a lecture on the challenges of ensuring AI systems follow instructions, using the metaphor of Asimov’s Three Laws of Robotics. He begins with a brief history of AI, from symbolic reasoning to deep learning, highlighting the AI winter and the recent transformer-based breakthroughs. The core of the talk is a hands-on demonstration using Comet’s OPIC platform to test LLM robustness against jailbreaking attempts. He shows how to create a system prompt with a secret password and then attempts to extract it through various prompts, illustrating the ease with which LLMs can be manipulated. He introduces the concept of systematic evaluation using datasets, metrics, and experiments, and demonstrates how to compare different models (GPT-4o Mini, GPT-4o, GPT-5, Gemini 2.5 Flash) on a jailbreak dataset. He also discusses the use of LLM-as-a-judge for more nuanced evaluation. The talk concludes with a case study of a real-world incident where an LLM provided harmful advice, emphasizing the importance of rigorous testing and the need for continuous improvement in AI safety.

175 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the practical aspects of LLM evaluation and safety. The live demonstrations are effective in illustrating the vulnerabilities of current models. The argumentation is persuasive, drawing on personal experience and a real-world case study to underscore the importance of systematic testing. However, the talk is more of an expert opinion and tutorial than a rigorous scientific study, and the evidence is largely anecdotal.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically grounded, referencing the history of AI and the transformer architecture. The sources cited are primarily the course website and the OPIC platform, which are relevant. The title is somewhat misleading as it does not directly address Asimov’s laws but uses them as a metaphor for AI safety. The content is well-structured and technically accurate, but the lack of formal citations and reliance on personal experience slightly reduces its scientific rigor.

157 words

Title / Content Match

The title is somewhat metaphorical, referencing Asimov's laws, but the content focuses on practical AI safety and evaluation, which is only loosely connected.

Quality & Reliability

8/10

The talk is given by a researcher with deep expertise in AI, presenting practical tools and methodologies for LLM evaluation. It includes live demonstrations and references to real incidents, but relies heavily on personal experience and anecdotal evidence rather than formal studies.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk provides a practical framework for evaluating LLM robustness against jailbreaking, using a systematic approach with datasets and metrics. It highlights the importance of continuous testing and optimization of prompts and models. The live demonstration with audience participation makes the concepts tangible.

Pour aller plus loin :

  • Prompt injection — Relevant to jailbreaking techniques.
  • AI alignment — Discusses the broader challenge of ensuring AI systems act in accordance with human intentions.
  • LLM-as-a-judge — A paper on using LLMs for evaluation.

81 words

Radar Profile

The radar profile shows high scores in information quantity and quality, with a moderate technical level and reliability. This indicates a well-informed talk with practical insights, but not deeply technical or formally rigorous.

Reliability 7/10