arXiv:2605.15205cs.AI2026-05ACL

测试了大模型心智理论能力对人机互动的实际效果,发现静态测试提升不等于真实交互变好。

Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations

论文配图:Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations
图 1 · 摘自论文原文
  • 设计新评测范式,从第三人称问答转为动态第一人称交互
  • 四种心智理论增强方法在真实任务中表现参差,与静态测试结果不一致
  • 强调必须用互动评估来研发真正懂人的下一代智能体

提升大语言模型的心智理论(ToM)能力对实现有效的人机社会互动至关重要。然而,现有基准大多通过阅读故事、第三人称选择题等静态方式衡量ToM,忽略了人机交互中第一人称、动态、开放性的特点。为直接检验ToM改进技术对人机互动的实际益处,本文首次提出交互式ToM评测范式,包含视角与指标双重转变。随后,基于该范式,对四种代表性ToM增强技术进行了系统研究,涵盖四个真实数据集和用户实验,覆盖目标导向任务(如编程、数学)与体验导向任务(如心理咨询)。结果表明,静态基准上的性能提升并不总能转化为动态人机交互中的实际改善。本研究揭示了当前评测体系的局限性,强调开发下一代具备社会意识的智能体需依赖交互式评估。

原文摘要 · Abstract (English)

Improving the Theory of Mind (ToM) capability of Large Language Models (LLMs) is crucial for effective social interactions between these AI models and humans. However, the existing benchmarks often measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic, and open-ended nature of human-AI (HAI) interactions. To directly examine how ToM improvement techniques benefit HAI interactions, we first proposed the new paradigm of interactive ToM evaluation with both perspective and metric shifts. Next, following the paradigm, we conducted a systematic study of four representative ToM enhancement techniques using both four real-world datasets and a user study, covering both goal-oriented tasks (e.g., coding, math) and experience-oriented tasks (e.g., counseling). Our findings reveal that improvements on static benchmarks do not always translate to better performance in dynamic HAI interactions. This paper offers critical insights into ToM evaluation, showing the necessity of interaction-based assessments in developing next-generation, socially aware LLMs for HAI symbiosis.

心智理论人机交互评测范式大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。