The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior

The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior

🎙 Angelina Wang 👥 75K 📅 January 26, 2026 ⏱ 31 min 👁 389 📄 expert opinion 🧭 2026-08-06
Available in: English (current) Français

Keywords

offline evaluationpersonalizationfield studysock puppetsMMLU

Summary

Angelina Wang’s talk at the Simons Institute addresses the inadequacy of standard offline evaluations for large language models (LLMs) in capturing real-world behavior influenced by personalization. She defines personalization as the automatic customization of responses to individual users, contrasting it with prompt engineering. She highlights that implementations vary across chatbots (ChatGPT, Meta AI, Gemini) and are opaque, leading to unpredictable behaviors. To study this, she conducted a field study with 800 real users (400 for ChatGPT, 400 for Gemini) who posed benchmark and recommendation questions to their personalized chatbots. Comparing these responses to offline API queries, she found greater diversity and sometimes unseen answer choices in field responses, even for objective science questions from MMLU. She also introduced ‘sock puppets’—simulated users with different histories—to scale evaluations, showing they increase diversity but still underrepresent real-world variation. The implications are that offline benchmark scores may not reflect individual user experiences, and models can exhibit different capabilities across users. She concludes by emphasizing the need to evaluate AI systems in human interaction contexts rather than decontextualized outputs.

174 words

Critical Evaluation

The talk provides a compelling and well-structured argument for reconsidering how LLMs are evaluated. The speaker presents empirical evidence from a field study involving 800 users, which is a significant strength, as it moves beyond theoretical concerns to real-world data. The methodology is clearly explained: participants were recruited via Prolific, with balanced demographics (Black men, Black women, White men, White women), and asked to copy-paste questions into their own ChatGPT or Gemini interfaces. This approach captures the personalized context that offline evaluations miss. The comparison between offline and field responses reveals notable differences: field responses are more diverse, and for objective questions like MMLU, users sometimes receive answer choices never seen in offline runs. This is a critical finding, as it suggests that benchmark scores may not reflect the actual performance experienced by users. The introduction of ‘sock puppets’ as a scalable alternative is innovative, though the speaker acknowledges they do not fully replicate real-world diversity. The talk also touches on social implications, such as differential capabilities across users, which could exacerbate inequalities. However, the presentation is based on a limited set of questions (13 in the field study, 514 MMLU questions for sock puppets), and the results may not generalize broadly. The speaker does not delve into potential confounding factors, such as user behavior or interface differences, that might influence responses. Additionally, the talk is a research presentation, not a peer-reviewed publication, so the findings should be considered preliminary. The argumentation is logical and well-supported, but the lack of detailed statistical analysis in the talk leaves some questions unanswered. Overall, the talk is valuable for highlighting a critical gap in LLM evaluation and offers a practical direction for improvement.

280 words

Title / Content Match

The title accurately reflects the core argument: offline evaluations are inadequate because they ignore personalization, which significantly alters model behavior.

Quality & Reliability

8/10

The talk presents empirical evidence from a field study with 800 users and compares offline evaluations with personalized interactions. The methodology is clearly described, and the results are statistically grounded. However, the talk is a presentation of ongoing research, and some details are not fully peer-reviewed yet.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

This talk contributes original empirical evidence on how personalization affects LLM behavior in real-world settings, highlighting the inadequacy of offline evaluations. It introduces a field study methodology and sock puppets as scalable proxies, showing that offline benchmarks may not reflect user experiences. The findings have implications for AI evaluation standards and fairness.

Pour aller plus loin :

  • MMLU benchmark — The benchmark used in the study to evaluate model capabilities.
  • WildChat dataset — A dataset of real user-chat interactions used for sock puppet histories.
  • Retrieval-Augmented Generation (RAG) — A technique used in one sock puppet variant to retrieve relevant history.

100 words

Radar Profile

The radar profile shows high scores in quantity and quality of information, with slightly lower technical depth and reliability, reflecting the empirical yet preliminary nature of the research.

Reliability 8/10