用真人数据模拟人类,低成本测试人机协作能力。
Ad-Hoc Human-AI Coordination Challenge
- 构建真人行为代理,替代真实人类进行评估
- 提供3079场游戏数据,限制数据量以促发高效学习
- 支持两人和三人场景,适合研究高效人机协同方法
实现人工智能与人类的无缝协作对现实应用至关重要,但仍是重大开放挑战。汉诺比(Hanabi)是一种合作类卡牌游戏,具有不完全信息、有限沟通、心智理论需求和协同行动等特点,是检验人机协作的理想实验平台。然而,其在人机交互中的应用受限于人力评估成本高且难以复现。为此,本文提出即兴人机协作挑战(AH2AC2),通过大规模人类数据集训练出可替代真人、成本低、可复现的‘人类代理’,用于评估。为鼓励数据效率高的方法发展,我们开源了3,079场游戏数据,刻意限制可用人类游戏数据量。本文提供了两人和三人汉诺比场景的基线结果。为确保评估公平性,代理模型通过受控系统托管,不对外公开。代码已发布于https://github.com/FLAIROx/ah2ac2。
原文摘要 · Abstract (English)
Achieving seamless coordination between AI agents and humans is crucial for real-world applications, yet it remains a significant open challenge. Hanabi is a cooperative card game featuring imperfect information, constrained communication, theory of mind requirements, and coordinated action -- making it an ideal testbed for human-AI coordination. However, its use for human-AI interaction has been limited by the challenges of human evaluation. In this work, we introduce the Ad-Hoc Human-AI Coordination Challenge (AH2AC2) to overcome the constraints of costly and difficult-to-reproduce human evaluations. We develop \textit{human proxy agents} on a large-scale human dataset that serve as robust, cheap, and reproducible human-like evaluation partners in AH2AC2. To encourage the development of data-efficient methods, we open-source a dataset of 3,079 games, deliberately limiting the amount of available human gameplay data. We present baseline results for both two- and three- player Hanabi scenarios. To ensure fair evaluation, we host the proxy agents through a controlled evaluation system rather than releasing them publicly. The code is available at \href{https://github.com/FLAIROx/ah2ac2}{https://github.com/FLAIROx/ah2ac2}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。