让大模型当汉诺比牌局队友,探索协作推理的极限
Sparks of Cooperative Reasoning: LLMs as Strategic Hanabi Agents
- 用工作记忆和贝叶斯推理增强提示,提升模型协作能力
- 最强模型平均得分15分,仍低于人类20+的稳定表现
- 数据集和训练方法可迁移至其他协作任务,效果显著
在信息不完全条件下进行协作推理对人类和多智能体系统均具挑战性。汉诺比牌戏体现了这一难题,需具备心理理论和策略沟通能力。我们对17个最先进的大模型智能体在2至5人游戏中进行了基准测试,研究了从4B到600B+不同规模模型下上下文设计的影响:从仅含明确卡牌信息的最小提示(Watson设定),到基于程序化、贝叶斯动机的推理解析(Sherlock),再到通过工作记忆实现多轮状态追踪(Mycroft设定)。结果表明:(1) 模型能维持内部工作记忆以追踪状态;(2) 不同模型间的跨游戏表现随模型能力平滑变化。在Sherlock设定下,最强推理模型平均得分超过15分,但仍落后于经验丰富的玩家及专用汉诺比智能体,二者均稳定高于20分。我们发布了首个公开的汉诺比数据集,包含标注轨迹与动作价值:(1) HanabiLogs,含1,520场完整游戏日志,用于指令微调;(2) HanabiRewards,含560场游戏,提供所有候选动作的细粒度移动级价值标注。在我们的数据集上,对4B开源模型Qwen3-Instruct进行监督和强化学习微调,使协作汉诺比表现分别提升21%和156%,性能接近强效专有推理模型(o4-mini),领先非推理模型(GPT-4.1)达52%。该汉诺比奖励强化学习微调模型还泛化至其他任务:在协作猜谜基准上提升11%,在EventQA时间推理上提升6.4分,在IFBench指令遵循上提升1.7 Pass@10,达到AIME 2025数学推理的Pass@10水平。
原文摘要 · Abstract (English)
Cooperative reasoning under incomplete information remains challenging for both humans and multi-agent systems. The card game Hanabi embodies this challenge, requiring theory-of-mind reasoning and strategic communication. We benchmark 17 state-of-the-art LLM agents in 2-5 player games and study the impact of context engineering across model scales (4B to 600B+) to understand persistent coordination failures and robustness to scaffolding: from a minimal prompt with only explicit card details (Watson setting), to scaffolding with programmatic, Bayesian-motivated deductions (Sherlock), to multi-turn state tracking via working memory (Mycroft setting). We show that (1) agents can maintain an internal working memory for state tracking and (2) cross-play performance between different LLMs smoothly interpolates with model strength. In the Sherlock setting, the strongest reasoning models exceed 15 points on average across player counts, yet still trail experienced humans and specialist Hanabi agents, both consistently scoring above 20. We release the first public Hanabi datasets with annotated trajectories and move utilities: (1) HanabiLogs, containing 1,520 full game logs for instruction tuning, and (2) HanabiRewards, containing 560 games with dense move-level value annotations for all candidate moves. Supervised and RL finetuning of a 4B open-weight model (Qwen3-Instruct) on our datasets improves cooperative Hanabi play by 21% and 156% respectively, bringing performance to within ~3 points of a strong proprietary reasoning model (o4-mini) and surpassing the best non-reasoning model (GPT-4.1) by 52%. The HanabiRewards RL-finetuned model further generalizes beyond Hanabi, improving performance on a cooperative group-guessing benchmark by 11%, temporal reasoning on EventQA by 6.4 points, instruction-following on IFBench by 1.7 Pass@10, and matching AIME 2025 mathematical reasoning Pass@10.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。