Sam Eisenstat - Concepts, information, and objectivity - IPAM at UCLA

Sam Eisenstat - Concepts, information, and objectivity - IPAM at UCLA

🎙 Sam Eisenstat 👥 42K 📅 September 2, 2026 ⏱ 44 min 👁 4 📄 original study 🧭 2026-09-02
Available in: English (current) Français

Keywords

latent variable modelinformation theoryuniqueness theoremconceptinterpretability

Summary

Sam Eisenstat presents a theoretical framework for understanding how different agents (humans, AI systems) can share concepts and representations. He motivates the problem with examples like language learning and interpretability methods (e.g., Golden Gate Claude). The core contribution is a mathematical model using latent variables: observed variables (e.g., texts) are deterministic functions of latent variables, and the goal is to show that under certain conditions, such latent variable models are approximately unique (up to isomorphism). He defines a ‘contribution relation’ and introduces conditions such as a Markov condition (approximate independence of latent variables) and a reconstruction condition (observing enough descendants allows inference of the latent variable). The main theorem states that two latent variable models satisfying these conditions admit an approximate bijection between their latent variables, with structural and informational correspondence. He discusses examples, including a coin with unknown bias, and relates the framework to Bayesian networks. The talk concludes by discussing implications for interpretability and machine learning, emphasizing the need for a theoretical grounding of empirical observations.

168 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides a novel theoretical contribution by formalizing the notion of shared concepts through latent variable models and information-theoretic uniqueness. The argumentation is rigorous: the speaker carefully defines the model, states conditions, and sketches the proof of the main theorem. He also addresses potential objections and clarifies technical points in response to audience questions. The value lies in offering a principled framework that could complement empirical interpretability research, though the practical applicability to real neural networks remains an open question.

Scientific Rigor, Source Quality, Title Accuracy

The talk is based on a paper by the speaker and subsequent work, though no specific references are given in the video. The presentation is mathematically rigorous, with clear definitions and a theorem. The title accurately reflects the content. The description provides a link to the IPAM workshop page, which is the primary source. No external sources are cited in the video itself.

159 words

Title / Content Match

The title accurately reflects the content: the talk introduces a theoretical model for concepts using latent variables and information theory, aiming to establish a form of objectivity (uniqueness) in representation.

Quality & Reliability

8/10

The talk presents a formal mathematical framework with a theorem and proof sketch, grounded in probability theory and information theory. The approach is rigorous, with clear definitions and conditions, though the presentation is at a research seminar level and the theorem's applicability to real-world interpretability remains to be validated empirically.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk offers a novel theoretical framework for understanding concept sharing and objectivity in representations, using latent variable models and information theory. It provides a formal uniqueness theorem that could underpin interpretability research. The approach is original in its combination of ideas from Bayesian networks, algorithmic information theory, and statistical learning.

Pour aller plus loin :

115 words

Radar Profile

The radar profile shows high scores in technical level and information quality, reflecting the formal mathematical nature of the talk. The quantity of information is moderate, as the talk is focused on a specific theoretical result. The global reliability is high due to the rigorous presentation, but the practical applicability remains uncertain.

Reliability 8/10