arXiv:2504.10839cs.HCcs.AI2025-04中稿 · the HEAL@CHI 2025 …被引 21

重新审视大模型心理理论评测,强调用户中心的交互视角。

Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective

  • 从人机交互角度重构大模型心智理论评测标准
  • 指出传统人类心理理论任务在评测大模型时存在固有缺陷
  • 适合关注大模型社会智能与用户体验的研究者

近年来,研究者将专为人类设计的心理理论(ToM)任务用于评估大模型的社会智能能力。然而,这一方法存在诸多局限。基于心理学与人工智能领域的文献,我们总结了理论、方法和评估层面的问题,指出原始人类ToM任务中的某些固有问题在应用于大模型评测时不仅持续存在,反而被加剧。从人机交互(HCI)视角出发,我们主张以更动态、互动的方式重新定义和构建大模型的ToM评测体系,纳入用户偏好、需求与使用体验。最后,我们展望了该方向的潜在机遇与挑战。

原文摘要 · Abstract (English)

The last couple of years have witnessed emerging research that appropriates Theory-of-Mind (ToM) tasks designed for humans to benchmark LLM's ToM capabilities as an indication of LLM's social intelligence. However, this approach has a number of limitations. Drawing on existing psychology and AI literature, we summarize the theoretical, methodological, and evaluation limitations by pointing out that certain issues are inherently present in the original ToM tasks used to evaluate human's ToM, which continues to persist and exacerbated when appropriated to benchmark LLM's ToM. Taking a human-computer interaction (HCI) perspective, these limitations prompt us to rethink the definition and criteria of ToM in ToM benchmarks in a more dynamic, interactional approach that accounts for user preferences, needs, and experiences with LLMs in such evaluations. We conclude by outlining potential opportunities and challenges towards this direction.

心理理论大模型评测人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。