arXiv:2501.08838cs.CLcs.AI2025-01AAAI被引 24

用角色扮演让大模型自述心理状态,构建更真实的心智理论评测数据集。

ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind

  • 通过让大模型在对话前自述想法,捕捉五类心理状态。
  • 生成5400个问题,包含虚假信念和15种人格设定。
  • 适合评估大模型对复杂心理状态的理解能力。

现有心智理论(ToM)评测在三个方面与真实场景脱节:1)仅评估有限的心理状态如信念;2)未充分探索错误信念;3)忽略角色的多样人格特征。为解决这些问题,我们提出ToMATO,一种基于对话的多选题评测基准。ToMATO通过大模型间存在信息不对称的角色扮演对话生成。采用提示法要求角色扮演的大模型在每次发言前自述其思想,从而捕获第一、第二层心理状态,涵盖信念、意图、欲望、情绪和知识五类。这些自述内容作为问题答案,用于评估对话中角色的心理状态。通过隐藏思想制造信息不对称,诱导生成关于各类心理状态的错误信念。为角色分配不同人格特征,进一步丰富话语与思想表达。ToMATO包含5.4k个问题、753段对话和15种人格模式。分析表明,该构建方法因角色间信息不对称频繁生成错误信念,并有效反映多样性人格。我们在ToMATO上评估了九个大模型,发现即使GPT-4o mini也落后于人类表现,尤其在理解错误信念方面,且对不同人格缺乏鲁棒性。

原文摘要 · Abstract (English)

Existing Theory of Mind (ToM) benchmarks diverge from real-world scenarios in three aspects: 1) they assess a limited range of mental states such as beliefs, 2) false beliefs are not comprehensively explored, and 3) the diverse personality traits of characters are overlooked. To address these challenges, we introduce ToMATO, a new ToM benchmark formulated as multiple-choice QA over conversations. ToMATO is generated via LLM-LLM conversations featuring information asymmetry. By employing a prompting method that requires role-playing LLMs to verbalize their thoughts before each utterance, we capture both first- and second-order mental states across five categories: belief, intention, desire, emotion, and knowledge. These verbalized thoughts serve as answers to questions designed to assess the mental states of characters within conversations. Furthermore, the information asymmetry introduced by hiding thoughts from others induces the generation of false beliefs about various mental states. Assigning distinct personality traits to LLMs further diversifies both utterances and thoughts. ToMATO consists of 5.4k questions, 753 conversations, and 15 personality trait patterns. Our analysis shows that this dataset construction approach frequently generates false beliefs due to the information asymmetry between role-playing LLMs, and effectively reflects diverse personalities. We evaluate nine LLMs on ToMATO and find that even GPT-4o mini lags behind human performance, especially in understanding false beliefs, and lacks robustness to various personality traits.

心智理论角色扮演虚假信念大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。