Making AI Systems Work for Imperfect Humans

Making AI Systems Work for Imperfect Humans

🎙 Sherry Tongshuang Wu 👥 843 📅 August 12, 2026 ⏱ 57 min 👁 8 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

imperfect usersinteraction trajectorieshuman effortutilityAI agents

Summary

Sherry Wu’s talk addresses the challenge of designing AI systems that effectively support users who are not perfect oracles of their own intentions. She begins with a motivating example from her teaching experience, where students with varying backgrounds used AI assistants like Gemini in Colab. Analysis revealed that success depended not on initial expertise but on interaction style: some students iteratively refined their tasks with the AI, while others struggled to communicate their needs. Wu generalizes this to a broader observation that current AI assumes perfect users with fixed, well-defined goals, which is unrealistic. She proposes a framework called ‘collaborative effort scaling’ to evaluate human-AI interaction processes along two dimensions: stability (how utility scales with human effort) and usability (how much effort users are willing to invest before dropping out). She illustrates this with simulation data from Collaborative Gym, showing that different models exhibit different scaling patterns. The talk concludes by discussing two approaches to computationally operationalize the user for better intervention: one focusing on training models to better guess user needs, and another on training humans to better express themselves. Overall, the talk provides a thoughtful perspective on evaluating and improving human-AI collaboration.

194 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk offers valuable insights into the often-overlooked aspect of human imperfection in AI interaction. Wu’s argument is well-structured, moving from a concrete teaching example to a general framework. She effectively uses examples to illustrate key points, such as the column name error and the pie chart iteration. The framework of collaborative effort scaling is a novel contribution that could influence how AI systems are evaluated. The argumentation is solid, though some claims are based on anecdotal evidence from a single classroom setting. The discussion of simulation results adds empirical weight, but the talk would benefit from more details on the methodology and metrics used.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor by grounding the framework in prior work, such as the 1994 paper on mixed-initiative interfaces. Wu references her own published research and the Collaborative Gym simulation environment, which adds credibility. However, specific citations are not always provided in the talk, and the description only includes a personal website. The title accurately reflects the content, and the talk stays on topic. The lack of detailed citations in the talk itself is a minor weakness, but the overall rigor is high.

204 words

Title / Content Match

The title accurately reflects the content, focusing on designing AI systems for imperfect human users.

Quality & Reliability

8/10

The talk is given by a recognized expert in HCI and NLP, presenting a coherent framework based on empirical observations from classroom and simulation studies. The claims are supported by examples and references to prior work, though some details are anecdotal.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk introduces a novel evaluation framework, ‘collaborative effort scaling’, that shifts focus from final output to the interaction process itself. This is a valuable contribution to HCI and AI evaluation, as it acknowledges the dynamic nature of human goals and the importance of sustaining user engagement. The framework provides concrete dimensions (stability and usability) that can be measured and used to compare AI agents. The talk also highlights the need for interventions that either improve the model’s ability to infer user needs or enhance the user’s ability to communicate, opening new research directions.

Pour aller plus loin :

  • Mixed-Initiative User Interfaces — The 1994 paper by Horvitz is foundational for understanding mixed-initiative interaction, which is central to the talk’s theme.
  • Human-in-the-loop — This concept is directly relevant to the talk’s focus on human-AI collaboration and the need for iterative interaction.
  • Collaborative Gym — The simulation environment used in the talk to study human-AI collaboration; provides a platform for further research.

161 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a moderate technical level. This reflects a talk that is rich in content and well-supported, but not overly technical, making it accessible to a broad audience.

Reliability 8/10