解决多智能体协作中贡献分配难题,按合作层级精准评估每个智能体的贡献。
Multi-level Advantage Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- 提出多层级优势函数,通过反事实推理区分不同合作层次的贡献。
- 在星际争霸1/2任务中表现优于现有方法,显著提升复杂协作场景下的性能。
- 适合研究多智能体强化学习、分布式系统协作的开发者与研究人员。
协作式多智能体强化学习(MARL)旨在协调多个智能体共同实现目标。其核心挑战之一是信用分配问题,即评估每个智能体对共享奖励的贡献。由于任务多样性,智能体可能以不同方式协作,且奖励常由重叠的智能体子集获得。本文将信用分配层级定义为协同获取奖励的智能体数量,处理多种层级共存的情况。提出多层级优势形式化方法,通过显式反事实推理推断不同层级的信用分配。所提方法MACA通过整合针对个体、联合及关联动作的优势函数,捕捉多层级贡献。利用基于注意力的框架,识别智能体间的关联关系,并构建多层级优势以指导策略学习。在具有挑战性的Starcraft v1&v2任务上的综合实验表明,MACA性能优越,验证了其在复杂信用分配场景中的有效性。
原文摘要 · Abstract (English)
Cooperative multi-agent reinforcement learning (MARL) aims to coordinate multiple agents to achieve a common goal. A key challenge in MARL is credit assignment, which involves assessing each agent's contribution to the shared reward. Given the diversity of tasks, agents may perform different types of coordination, with rewards attributed to diverse and often overlapping agent subsets. In this work, we formalize the credit assignment level as the number of agents cooperating to obtain a reward, and address scenarios with multiple coexisting levels. We introduce a multi-level advantage formulation that performs explicit counterfactual reasoning to infer credits across distinct levels. Our method, Multi-level Advantage Credit Assignment (MACA), captures agent contributions at multiple levels by integrating advantage functions that reason about individual, joint, and correlated actions. Utilizing an attention-based framework, MACA identifies correlated agent relationships and constructs multi-level advantages to guide policy learning. Comprehensive experiments on challenging Starcraft v1\&v2 tasks demonstrate MACA's superior performance, underscoring its efficacy in complex credit assignment scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。