arXiv:2502.00345cs.LGcs.AI2025-02

提出新型协作多智能体任务挑战,检验分工与合作能力。

CTC: The Composite Task Challenge for Cooperative Multi-Agent Reinforcement Learning

  • 设计需明确分工与协作才能完成的任务集
  • 现有9种主流方法在所有任务上胜率均为0%
  • 提供可解基准方案,凸显任务挑战性

分工(DOL)在现实应用中对提升协作至关重要,许多协作多智能体强化学习(MARL)方法已引入DOL机制。然而,缺乏专门用于评估和推动DOL与协作的基准任务,限制了这些机制在实际中的发展。为填补这一空白,我们提出复合任务挑战(CTC),一套明确要求分工与协作才能成功完成的任务集。其设计基于两个核心原则:1)分工是任务成功的必要条件;2)任意子任务失败将导致整体任务失败。前者强调分工必要性,后者强化协作重要性,两者共同构成成功关键。我们在新提出的CTC任务上评估了九种代表性协作MARL方法,结果均显示测试胜率为0%,凸显任务难度与当前方法局限性。为此,我们引入一个引导性解决方案,在所有任务上实现非零胜率,证明任务可解。但该方案表现仍不理想,进一步验证了CTC作为推进协作MARL研究的高挑战性、有意义基准的价值。

原文摘要 · Abstract (English)

The critical role of division of labor (DOL) in enhancing cooperation is well-recognized in real-world applications. Consequently, many cooperative multi-agent reinforcement learning (MARL) methods have incorporated DOL mechanisms to improve cooperation among agents. However, the lack of benchmark tasks specifically designed to evaluate and promote DOL and cooperation has limited the effective development and deployment of such mechanisms in cooperative MARL. This gap between current cooperative MARL methods and practical applications underscores the need for evaluation tasks that explicitly require DOL and cooperation. To address this gap, we propose the Composite Tasks Challenge (CTC), a suite of tasks explicitly designed to require both DOL and cooperation for successful task completion. The CTC tasks are constructed based on two core design principles: 1) DOL is a necessary condition for task success; 2) Failure in any atomic subtask results in failure of the overall task. The first principle emphasizes the necessity of DOL, while the second enforces the importance of cooperation, making both components essential for success in CTC tasks. We evaluate nine representative cooperative MARL methods on the proposed CTC tasks. Experimental results show that all methods consistently achieve zero test winning rates across all CTC tasks, highlighting the challenge of CTC tasks and the limitations of current methods. To facilitate future research, we also introduce a guiding solution that achieves non-zero test winning rates on all tasks, thereby demonstrating the solvability of the CTC tasks. However, the performance of this guiding solution remains suboptimal, further underscoring the value of CTC tasks as a challenging and meaningful testbed for advancing cooperative MARL research.

多智能体强化学习任务设计协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。