Dana Arad: SAEs for Content Control

Dana Arad: SAEs for Content Control

🎙 Dana Arad 👥 843 📅 August 12, 2026 ⏱ 60 min 👁 25 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

SAEsteeringinterpretabilityunlearningfeatures

Summary

Dana Arad presents two recent papers on using Sparse Autoencoders (SAEs) for content control in language models. She first introduces the concept of polysemanticity and how SAEs decompose latent spaces into interpretable features. She then discusses the challenge of feature selection for steering, highlighting that activation-based selection often fails. The first paper introduces a taxonomy distinguishing input features (which process input) from output features (which causally affect output), and proposes scores to quantify these roles. They show that output features emerge in later layers, aligning with model phase transitions. The second paper presents a method for persistent unlearning using SAE features, enabling fine-grained control. The talk emphasizes the potential of SAEs for interpretable control while acknowledging the need for deeper understanding. The presentation includes interactive Q&A sessions and references to prior work, such as the AxBench benchmark and Anthropic’s gender bias feature.

142 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the practical use of SAEs for steering and unlearning, addressing a critical gap in the literature. The argumentation is solid, supported by empirical results and clear methodology. The speaker effectively motivates the problem by showing failures of naive feature selection and then presents a systematic approach to identify output features. The discussion of layer-wise trends adds depth, and the open questions encourage further research. The talk is well-structured and the speaker is responsive to audience questions, clarifying technical details.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor by referencing prior work (e.g., Anthropic’s gender bias feature, AxBench benchmark) and presenting original research with clear definitions and metrics. The sources cited are relevant and credible. The title accurately reflects the content, focusing on SAEs for content control. The talk is a presentation of ongoing research, so some claims are not yet peer-reviewed, but the methodology is sound. The speaker also acknowledges limitations and open questions, enhancing credibility.

174 words

Title / Content Match

The title accurately reflects the content, focusing on using SAEs for content control.

Quality & Reliability

8/10

The talk presents original research from two papers, with clear methodology and empirical results. The speaker is a PhD candidate with relevant expertise. However, the talk is a presentation, not a peer-reviewed publication, and some claims rely on unpublished work.

Key Moments

Cited Sources

  • SAEs for Content Control (paper 1) — Presented in the talk, distinguishing input and output features.
  • Persistent unlearning with SAEs (paper 2) — Presented in the talk, method for fine-grained unlearning.
  • Steering LLMs with SAEs (Anthropic) — Referenced for gender bias feature example.
  • AxBench benchmark — Referenced for evaluating steering methods.

Concurring Sources

Dissenting Sources

  • Steering LLMs with SAEs (Stanford) — Showed that simple baselines outperform SAEs, contradicting the claim that SAEs are effective for steering.

Contribution & Novelties

The talk presents original contributions: a taxonomy of SAE features (input vs output) and a method for persistent unlearning using SAE features. This addresses a critical gap in feature selection for steering, showing that activation-based selection is insufficient. The layer-wise analysis provides insights into feature roles across model depth. The work is novel and has practical implications for interpretable control of LLMs.

Pour aller plus loin :

95 words

Radar Profile

The radar profile shows high scores in information quantity, quality, technical level, and reliability, indicating a well-rounded and credible presentation. The talk is technically deep and provides substantial information, with a strong foundation in prior work.

Reliability 8/10

💬 No comments were provided for analysis.