arXiv:2608.03644cs.AIcs.MA2026-08

检验零样本协作算法对实现差异的鲁棒性,发现标准评估已足够可靠。

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

论文配图:Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details
图 1 · 摘自论文原文
  • 设计跨实现对抗测试,模拟不同团队对同一规范的多样实现。
  • 针对热门算法Other-Play进行多实现评估,结果与传统方法高度一致。
  • 为零样本协作算法的评测提供更严谨的框架,适合算法开发者参考。

现实世界中部署的AI智能体需能与未见过的人类或其他智能体进行协作。零样本协作(ZSC)算法通过定义高层学习规则,使独立训练的智能体在测试时自动协同。然而,对ZSC算法的严格评估仍具挑战:理想情况下应使用多个独立实现版本,以反映不同团队对同一规范的理解差异。实践中,多数研究仅用单一实现、不同随机种子训练,少数还改变神经网络架构。这使得算法对规范模糊性和实现细节的鲁棒性存疑。本文首次系统评估该鲁棒性,提出跨实现对抗测试(cross-implementation cross-play),涵盖此前被证明影响多智能体强化学习性能的多种实现差异,并以Popular ZSC算法Other-Play为对象进行测试。结果表明,对于Other-Play,标准的ZSC评估可作为更全面跨实现评估的有效替代。

原文摘要 · Abstract (English)

AI agents deployed in real-world settings must be capable of coordinating with humans and other AI agents they have not encountered before. Zero-shot coordination (ZSC) algorithms aim to achieve this by specifying high-level learning rules such that independently engineered agents can coordinate with each other at test time. Rigorous evaluation of ZSC algorithms remains difficult: ideally, multiple independent implementations of each proposed algorithm must be used, reflecting the variation that arises when independent parties interpret and implement the same specification. In practice, however, ZSC algorithms have almost exclusively been evaluated using a single implementation trained across different random seeds, with only a handful of works additionally varying the neural network architecture. This leaves open questions about robustness to specification ambiguities and implementation details. In this work, we provide the first systematic evaluation of this robustness. We introduce a new evaluation scheme, cross-implementation cross-play, varying implementation details that prior work has shown to affect the performance of multi-agent reinforcement learning (MARL) algorithms, and we evaluate Other-Play, a popular ZSC algorithm, with this scheme. Our findings are encouraging and suggest that, for Other-Play, the standard ZSC evaluation is, in fact, a reasonable proxy for this more thorough cross-implementation evaluation.

零样本协作多智能体算法评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。