arXiv:2605.06557cs.MAcs.AI2026-05

提出新评估方法,揭示多智能体协作中的真实协调机制。

Coordination Matters: Evaluation of Cooperative Multi-Agent Reinforcement Learning

论文配图:Coordination Matters: Evaluation of Cooperative Multi-Agent Reinforcement Learning
图 1 · 摘自论文原文
  • 设计可控实验环境STAT,系统测试不同协作策略。
  • 相同回报下,协调方式差异显著,如任务重复分配、多样性等。
  • 适合研究多智能体协作效率与可扩展性的学者参考。

协作式多智能体强化学习(MARL)基准通常关注总回报、成功率或完成时间等聚合指标。然而,这些指标难以揭示智能体间的实际协调过程,尤其在智能体、任务及联合分配选择呈组合增长的场景中。本文提出一种注重协调过程的评估视角,补充传统回报指标。通过构建STAT——一个受控的、承诺约束的空间任务分配测试平台,系统地改变智能体数量、任务数量和环境规模,同时固定观测范围和任务规则。在不同中心化程度下评估六种代表性值函数型MARL方法。结果表明,相似的回报趋势可能对应截然不同的协调机制,包括冗余分配、分配多样性及任务完成效率的差异。在承诺约束的任务分配中,性能受规模影响不仅源于动作空间大小,还取决于分配压力、稀疏决策机会以及智能体间依赖下的冗余选择。研究支持将协调感知评估作为回报基准的必要补充。

原文摘要 · Abstract (English)

Cooperative multi-agent reinforcement learning (MARL) benchmarks commonly emphasize aggregate outcomes such as return, success rate, or completion time. While essential, these metrics often fail to reveal how agents coordinate, particularly in settings where agents, tasks, and joint assignment choices scale combinatorially. We propose a coordination-aware evaluation perspective that supplements return with process-level diagnostics. We instantiate this perspective using STAT, a controlled commitment-constrained spatial task-allocation testbed that systematically varies agents, tasks, and environment size while holding observation access and task rules fixed. We evaluate six representative value-based MARL methods across varying levels of centralization. Our results show that similar return trends can reflect distinct coordination mechanisms, including differences in redundant assignment, assignment diversity, and task-completion efficiency. We find that in commitment-constrained task allocation, performance under scale is shaped not only by nominal action-space size, but also by assignment pressure, sparse decision opportunities, and redundant choices among interdependent agents. Our findings motivate coordination-aware evaluation as a necessary complement to return-based benchmarking for cooperative MARL.

多智能体强化学习协作评估协调机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。