arXiv:2501.02221cs.AIcs.LG2025-01被引 2

通过角色多样性实现多智能体协作的泛化能力提升

CORD: Generalizable Cooperation via Role Diversity

  • 高层控制器通过最大化角色熵并施加约束,动态分配角色
  • 在多种任务中表现优于基线,尤其在未见合作者场景下泛化性更强
  • 适合需要灵活协作的复杂多智能体系统应用

合作式多智能体强化学习旨在训练能有效协同的智能体。然而,多数方法在训练智能体上过拟合,导致策略难以泛化到未见过的合作者,这制约了实际部署。部分方法虽尝试解决泛化问题,但需预先知道新队友的先验知识或预定义策略,限制了实际应用。为此,我们提出一种分层多智能体强化学习方法——CORD,通过角色多样性实现可泛化的协作。其高层控制器通过最大化角色熵并施加约束,为底层智能体分配角色。我们证明该约束目标可分解为角色的因果影响(确保合理分配)与角色异质性(生成一致且非冗余的角色集群)。在多种合作任务上的评估显示,CORD优于基线,尤其在泛化测试中表现突出。消融实验进一步验证了该约束目标在可泛化协作中的有效性。

原文摘要 · Abstract (English)

Cooperative multi-agent reinforcement learning (MARL) aims to develop agents that can collaborate effectively. However, most cooperative MARL methods overfit training agents, making learned policies not generalize well to unseen collaborators, which is a critical issue for real-world deployment. Some methods attempt to address the generalization problem but require prior knowledge or predefined policies of new teammates, limiting real-world applications. To this end, we propose a hierarchical MARL approach to enable generalizable cooperation via role diversity, namely CORD. CORD's high-level controller assigns roles to low-level agents by maximizing the role entropy with constraints. We show this constrained objective can be decomposed into causal influence in role that enables reasonable role assignment, and role heterogeneity that yields coherent, non-redundant role clusters. Evaluated on a variety of cooperative multi-agent tasks, CORD achieves better performance than baselines, especially in generalization tests. Ablation studies further demonstrate the efficacy of the constrained objective in generalizable cooperation.

多智能体强化学习角色分工泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。