Andrew Lee: Decomposing Query-Key Feature Interactions Using Contrastive Covariances

Andrew Lee: Decomposing Query-Key Feature Interactions Using Contrastive Covariances

🎙 Andrew Lee 👥 843 📅 August 12, 2026 ⏱ 60 min 👁 5 📄 original study 🧭 2026-08-16
Available in: English (current) Français

Keywords

attention headsquery-key spacecontrastive covarianceinterpretabilitylow-rank decomposition

Summary

Andrew Lee presents a method to interpret attention heads in Transformers by decomposing the query-key (QK) space into low-rank, human-interpretable components. He introduces a contrastive covariance approach that isolates the contribution of specific features (called ’tags’) to attention scores. The method is first validated on a toy task with synthetic data, where it successfully recovers the latent variables used to generate the data. Then, it is applied to large language models, identifying interpretable QK subspaces for categorical semantic features and binding features. The talk emphasizes that attention scores arise from alignment of keys and queries in these low-rank subspaces, and demonstrates how attention can be attributed to identified features. The approach offers a new tool for understanding why attention heads attend to particular tokens, moving beyond black-box attention patterns.

129 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides a valuable contribution to interpretability by addressing the under-explored question of why attention heads attend to specific tokens. The method is well-motivated and the argumentation is solid: the toy task clearly illustrates the mechanics, and the application to real LLMs demonstrates practical utility. The contrastive covariance approach is a novel way to decompose the QK space, and the results are convincing. The speaker also engages with audience questions, clarifying technical details and providing intuition.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous: the method is mathematically grounded, and the empirical validation is thorough. The speaker does not cite external sources during the talk, but the work is presumably based on a paper (not mentioned). The title accurately reflects the content. No comments were provided, so no analysis of public reception is possible.

147 words

Title / Content Match

The title accurately reflects the content: the talk focuses on decomposing query-key feature interactions using contrastive covariances.

Quality & Reliability

8/10

The talk presents a novel method with rigorous mathematical formulation, empirical validation on toy tasks and real LLMs, and is delivered by a recognized researcher in interpretability. The method is clearly explained and the results are convincing, though the talk is a presentation and not a peer-reviewed publication.

Key Moments

Contribution & Novelties

The talk introduces a novel method for interpreting attention heads by decomposing the QK space using contrastive covariances. This approach goes beyond existing interpretability tools like sparse autoencoders, which focus on activations, by directly analyzing the bilinear interaction space. The method provides a way to identify low-rank subspaces that correspond to human-interpretable features, enabling attribution of attention scores to specific features. This is a significant step towards understanding the internal mechanisms of Transformers.

Pour aller plus loin :

  • Sparse autoencoders — Related work on decomposing activations into interpretable features.
  • Attention is All You Need — Original Transformer paper introducing attention mechanisms.
  • Mechanistic Interpretability — Field of research focused on reverse-engineering neural networks.

112 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded presentation with substantial information, strong technical depth, and high reliability. The method is novel and clearly explained, making it a valuable contribution to the field.

Reliability 8/10