arXiv:2510.04980cs.AIcs.CL2025-10被引 7

用汉诺比游戏测试大模型的共情推理能力,发现理解对方意图比揣测对方想法更重要。

LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game

  • 用汉诺比合作游戏构建评估框架,自动衡量模型协作表现与共情推理能力。
  • 实验证明一阶共情(理解他人意图)与游戏成功显著正相关,优于二阶共情。
  • 适合研究多智能体协作、具身认知与大模型社会性推理的学者参考。

有效多智能体协作需要智能体推断他人行为背后的动机,这种能力源于心智理论(ToM)。尽管大型语言模型(LLMs)在逻辑推理方面表现优异,但其在动态协作场景中的动机推断能力仍待深入探索。本文提出 LLM-Hanabi,一个基于合作游戏汉诺比(Hanabi)的新基准,用于评估 LLM 的动机推断与心智理论能力。该框架包含自动化评估系统,可同时测量游戏表现与 ToM 水平。在多种模型上的实验表明,ToM 与游戏成功之间存在显著正相关。值得注意的是,一阶 ToM(解读他人意图)与性能的相关性高于二阶 ToM(预测他人对意图的理解)。结果表明,在有效人工智能协作中,准确理解伙伴动机的能力比高阶推理更为关键。因此,优先提升一阶 ToM 是增强未来模型协作能力的可行方向。

原文摘要 · Abstract (English)

Effective multi-agent collaboration requires agents to infer the rationale behind others' actions, a capability rooted in Theory-of-Mind (ToM). While recent Large Language Models (LLMs) excel at logical inference, their ability to infer rationale in dynamic, collaborative settings remains under-explored. This study introduces LLM-Hanabi, a novel benchmark that uses the cooperative game Hanabi to evaluate the rationale inference and ToM of LLMs. Our framework features an automated evaluation system that measures both game performance and ToM proficiency. Across a range of models, we find a significant positive correlation between ToM and in-game success. Notably, first-order ToM (interpreting others' intent) correlates more strongly with performance than second-order ToM (predicting others' interpretations). These findings highlight that for effective AI collaboration, the ability to accurately interpret a partner's rationale is more critical than higher-order reasoning. We conclude that prioritizing first-order ToM is a promising direction for enhancing the collaborative capabilities of future models.

多智能体心智理论协作推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。