Enric Boix Adsera

Enric Boix Adsera

🎙 Enric Boix Adsera 👥 4K 📅 May 3, 2026 ⏱ 28 min 👁 28 📄 expert opinion 🧭 2026-08-13
Available in: English (current) Français

Keywords

steeringalignmentlinear probingmulti-indexadversarial games

Summary

Enric Boix Adsera presents two research directions on AI alignment. First, he discusses steering models by editing internal activations, building on the linear representation hypothesis and extending it to a multi-index representation hypothesis. He introduces recursive feature machines (RFM) to train multi-index probes that outperform linear probes on tasks like hallucination and toxicity detection. He demonstrates that these probes can be used to steer models, e.g., to reduce refusal or increase honesty, by adding concept vectors to the residual stream. Second, he explores dynamic adaptation to stop misaligned agents in adversarial multi-agent settings, particularly cybersecurity. He shows that attackers can learn to evade defenders, and proposes a best-of-n self-distillation scheme where multiple defender copies flag attacks and the defender is trained on these flags, leading to a decrease in attack success rate. He provides a toy mathematical model and empirical results. The talk concludes with a Q&A session.

148 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk presents valuable insights into AI interpretability and alignment, with concrete examples and empirical results. The speaker argues for the multi-index representation hypothesis over the linear one, supported by experiments showing improved probe performance. The dynamic adaptation scheme for defenders is innovative and addresses a real problem, with a simple mathematical model and promising empirical results. However, the arguments are based on limited experiments and some informal setups, and the theoretical justification for the adaptation scheme is simplified.

Scientific Rigor, Source Quality, Title Accuracy

The speaker references several papers and sources, including the Alibaba incident, the linear representation hypothesis (Mikolov et al.), and representation engineering. He mentions his own papers and ongoing work. The sources are relevant and credible, but not all are formally cited with URLs. The title is just the speaker’s name, which is typical for seminar recordings and does not mislead. The content is rigorous for a research talk, but some claims are based on preliminary results.

170 words

Title / Content Match

The title is just the speaker's name, so it does not reflect the content, but this is typical for seminar recordings.

Quality & Reliability

7/10

Talk by a researcher presenting ongoing work, with references to papers and empirical results, but limited peer-reviewed validation and some informal experiments.

Key Moments

Cited Sources

  • Alibaba incident report — Mentioned as motivation for alignment issues.
  • Mikolov et al. word embeddings — Cited as origin of linear representation hypothesis.
  • Representation engineering paper — Referenced for steering method.

Concurring Sources

  • Anthropic research on monitoring — Mentioned as evidence that agents can evade monitors.
  • OpenAI research on monitoring — Mentioned as evidence that agents can evade monitors.

Contribution & Novelties

The talk presents novel approaches to AI alignment: extending linear probes to multi-index probes using recursive feature machines, and a dynamic adaptation scheme for defenders in adversarial games. These are original contributions, though some are in progress.

Pour aller plus loin :

70 words

Radar Profile

The radar profile shows high scores in quantity and technical level, with moderate quality and reliability, reflecting a research talk with substantial content but limited formal validation.

Reliability 6/10