arXiv:2502.06976cs.MAcs.AI2025-02AAAI被引 6

提出用相互依赖度衡量人机协作质量,发现高任务奖励不等于真合作。

Who is Helping Whom? Analyzing Inter-dependencies to Evaluate Cooperation in Human-AI Teaming

  • 用STRIPS形式化动作交互,量化人机对彼此行为的依赖程度。
  • 实验显示最优代理虽任务得分高,但人机间相互依赖极低。
  • 提醒研究者:任务奖励不能反映真实协作,适合评估团队合作的研究者参考。

当前人机协作与零样本协作方法仅以任务完成奖励为评价指标,忽视了双方实际如何协同。主观用户研究也难以揭示协作质量。本文关注训练好的智能体与人类配对时产生的合作行为,提出“建设性相互依赖”概念——衡量双方为达成共同目标对彼此行动的依赖程度,作为协作评价核心指标。基于STRIPS形式化,定义可量化行动间依赖关系的指标。在经典过厨房(Overcooked)场景中,将先进智能体HAT与学习的人类模型或真实人类参与者配对进行实验,评估任务奖励与协作表现。结果表明,尽管训练代理获得高任务奖励,但人机间的相互依赖程度极低,缺乏有效合作;且协作表现与任务奖励无必然关联,说明仅靠任务奖励无法可靠衡量团队内真实协作水平。

原文摘要 · Abstract (English)

State-of-the-art methods for Human-AI Teaming and Zero-shot Cooperation focus on task completion, i.e., task rewards, as the sole evaluation metric while being agnostic to how the two agents work with each other. Furthermore, subjective user studies only offer limited insight into the quality of cooperation existing within the team. Specifically, we are interested in understanding the cooperative behaviors arising within the team when trained agents are paired with humans -- a problem that has been overlooked by the existing literature. To formally address this problem, we propose the concept of constructive interdependence -- measuring how much agents rely on each other's actions to achieve the shared goal -- as a key metric for evaluating cooperation in human-agent teams. We interpret interdependence in terms of action interactions in a STRIPS formalism, and define metrics that allow us to assess the degree of reliance between the agents' actions. We pair state-of-the-art agents HAT with learned human models as well as human participants in a user study for the popular Overcooked domain, and evaluate the task reward and teaming performance for these human-agent teams. Our results demonstrate that although trained agents attain high task rewards, they fail to induce cooperative behavior, showing very low levels of interdependence across teams. Furthermore, our analysis reveals that teaming performance is not necessarily correlated with task reward, highlighting that task reward alone cannot reliably measure cooperation arising in a team.

人机协作合作评估相互依赖过厨房

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。