arXiv:2508.12480cs.AIcs.LG2025-08被引 5

新基准YLE挑战零样本协作,要求智能体动态跟踪信念与模糊提示。

The Yokai Learning Environment: Tracking Beliefs Over Space and Time

  • 设计可动态更新信念的多智能体环境,模拟真实协作中的信息不确定性。
  • 主流方法在YLE中表现显著下降,跨种子对战性能差距达30%以上。
  • 适合研究可信协作、信念建模与鲁棒零样本协调的学者使用。

在合作型人工智能中,与未知伙伴协作是一项核心挑战,常通过零样本协调(ZSC)来评估:即独立训练的智能体在配对时的表现。当前主流基准汉密尔顿学习环境(HLE)虽广泛使用,但近期方法已实现近乎完美的跨种子对战性能,限制了其对算法进步的追踪能力。为此,我们提出全新的开放源代码多智能体强化学习基准——妖物学习环境(YLE)。YLE要求有效协作需通过追踪移动牌面的信息、基于模糊提示推理,并依据推断的共享知识决定何时终止游戏,这些特性在HLE中并不存在;后者中信念固定于手牌位置,提示始终为真。我们评估了包括高熵IPPO、他者博弈与离信念学习在内的领先ZSC方法,它们在HLE中已达到近完美性能,但在YLE中却表现出持续的跨种子对战性能差距(SP-XP gap)、早期结束校准能力退化及交叉对战中信念表征弱化,表明无法与未知伙伴保持一致的内部模型。这说明在单一基准上取得进展并不意味着泛化能力,证明了YLE作为更具挑战性的新型ZSC基准的有效性。

原文摘要 · Abstract (English)

The ability to cooperate with unknown partners is a central challenge in cooperative AI and widely studied in the form of zero-shot coordination (ZSC), which evaluates an algorithm by measuring the performance of independently trained agents when paired. The Hanabi Learning Environment (HLE) has become the dominant benchmark for ZSC, but recent work has achieved near-perfect inter-seed cross-play performance, limiting its ability to track algorithmic progress. We introduce the Yokai Learning Environment (YLE) - an open-source multi-agent RL benchmark in which effective collaboration requires building common ground by tracking and updating beliefs over moving cards, reasoning under ambiguous hints, and deciding when to terminate the game based on inferred shared knowledge - features absent in the HLE, where beliefs are tied to hand slots and hints are truthful by rule. We evaluate the leading ZSC methods, including High-Entropy IPPO, Other-Play, and Off-Belief Learning, which achieve near-perfect inter-seed cross-play in the HLE, and show that in the YLE they exhibit persistent SP-XP gaps, degraded early-ending calibration, and weaker belief representations in cross-play, indicating failure to maintain consistent internal models with unseen partners. Methods that perform best in the HLE do not perform best in the YLE, indicating that progress measured on a single benchmark may not generalise. Together, these results establish YLE as a challenging new ZSC benchmark.

零样本协作多智能体信念建模强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。