Tianmin Shu: Scaling Model-based Theory of Mind for Socially Intelligent Embodied Partners (2026)

Tianmin Shu: Scaling Model-based Theory of Mind for Socially Intelligent Embodied Partners (2026)

🎙 Tianmin Shu 👥 843 📅 August 12, 2026 ⏱ 110 min 👁 28 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

Theory of MindEmbodied AIInverse PlanningHuman-Robot CollaborationMultimodal Reasoning

Summary

Tianmin Shu presents a research program on scaling model-based Theory of Mind (ToM) for socially intelligent embodied AI partners. He motivates the need for AI agents to understand human mental states (beliefs, goals) from behavior, beyond mere action recognition. He introduces the VirtualHome Social benchmark for multimodal ToM question answering, showing that humans perform well but current LLMs struggle, especially on goal inference. He proposes an automated approach (AUTOM) that combines LLMs for model construction and Bayesian inverse planning for inference, outperforming recent models on ToM benchmarks. He demonstrates the utility of ToM in a ‘Watch-and-Help’ human-robot collaboration task, where ToM-based planning improves assistance efficiency. He discusses proactive communication to align mental states and addresses challenges like noisy speech. The talk concludes with future directions on scaling ToM to real-world settings via self-supervised RL and grounding in 3D world models.

140 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into combining cognitive modeling with foundation models for ToM. The argumentation is solid, supported by empirical results from benchmarks and user studies. The speaker clearly explains the limitations of current LLMs and the benefits of explicit Bayesian inference. However, some claims about scalability and real-world deployment are forward-looking and not fully validated in the talk.

Scientific Rigor, Source Quality, Title Accuracy

The talk references several benchmarks and models (VirtualHome Social, AUTOM, Watch-and-Help) and mentions publications (ACL 2024 Outstanding Paper). The title accurately reflects the content. The speaker is a recognized researcher in the field, adding credibility. The talk does not include a formal literature review but provides sufficient context.

123 words

Title / Content Match

The title accurately reflects the content, focusing on scaling model-based Theory of Mind for embodied AI partners.

Quality & Reliability

8/10

The talk presents a coherent research program with published benchmarks and results, but as a conference talk it lacks full methodological details and peer-reviewed validation for all claims.

Key Moments

Cited Sources

  • VirtualHome Social — Benchmark for multimodal Theory of Mind reasoning in household environments.
  • AUTOM — Automated model construction for Theory of Mind using LLMs and Bayesian inference.
  • Watch-and-Help — Task for evaluating human-robot collaboration with Theory of Mind.

Concurring Sources

  • VirtualHome Social — Benchmark for multimodal Theory of Mind reasoning in household environments.
  • AUTOM — Automated model construction for Theory of Mind using LLMs and Bayesian inference.

Contribution & Novelties

The talk presents a novel approach combining LLMs with Bayesian inverse planning for scalable Theory of Mind, addressing limitations of pure LLM reasoning. It also introduces benchmarks and tasks for embodied ToM.

Pour aller plus loin :

59 words

Radar Profile

The radar profile shows high scores in information quantity and quality, with moderate technical depth and high reliability, indicating a well-rounded and credible presentation.

Reliability 8/10