新基准ToM-SSI测试模型在复杂社交环境中的心智推理能力。
ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions
- 设计多智能体、多模态的社交互动环境,支持四人组群交互。
- 现有模型在混合合作与阻挠场景中表现极差,准确率不足随机水平。
- 适合研究社会认知、多智能体协作的学者和开发者参考。
现有大模型理论心智(ToM)评估大多依赖萨莉-安妮测试变体,视角过于单一,忽视人类社交互动的复杂性。为弥补这一缺陷,我们提出ToM-SSI:一个专为具身社交互动环境设计的新基准。不同于仅限文本或双人交互的现有评测,ToM-SSI具备多模态特性,支持最多四个智能体在动态空间中进行交流与移动。该设计首次使研究者能够探索混合合作与阻挠情境,并并行推理多个智能体的心理状态,显著拓展了社会认知的评估范围。我们的评估显示,当前模型在这些新任务中表现严重受限,尤其在复杂情境下性能远低于随机水平,揭示出未来研究亟待填补的关键空白。
原文摘要 · Abstract (English)
Most existing Theory of Mind (ToM) benchmarks for foundation models rely on variations of the Sally-Anne test, offering only a very limited perspective on ToM and neglecting the complexity of human social interactions. To address this gap, we propose ToM-SSI: a new benchmark specifically designed to test ToM capabilities in environments rich with social interactions and spatial dynamics. While current ToM benchmarks are limited to text-only or dyadic interactions, ToM-SSI is multimodal and includes group interactions of up to four agents that communicate and move in situated environments. This unique design allows us to study, for the first time, mixed cooperative-obstructive settings and reasoning about multiple agents' mental state in parallel, thus capturing a wider range of social cognition than existing benchmarks. Our evaluations reveal that the current models' performance is still severely limited, especially in these new tasks, highlighting critical gaps for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。