通过构建竞争激励,让合作智能体更聪明多变。
Constructive Conflict-Driven Multi-Agent Reinforcement Learning for Strategic Diversity
- 用排名机制设计内在奖励,引入适度竞争。
- 在SMAC和GRF环境表现优于现有方法,策略更多样。
- 适合研究多智能体策略多样性与协作优化的读者。
近年来,多样性被证明是提升多智能体强化学习(MARL)效率的有效机制。然而,现有方法主要基于单个智能体特征设计策略,常忽略智能体间相互作用与影响。为此,我们提出竞争性多样性构建方法(CoDiCon),将竞争激励引入合作场景,促进策略交流并增强智能体间的战略多样性。受社会学研究启发——适度竞争与建设性冲突有助于群体决策,我们设计了一种基于排名特征的内在奖励机制,通过集中式奖励模块生成并分发不同奖励值,有效平衡竞争与合作。通过优化参数化中心奖励模块以最大化环境奖励,我们将约束型双层优化问题重构为与原任务目标一致的形式。在SMAC和GRF环境中的实验表明,CoDiCon取得优异性能,竞争性内在奖励有效促进了合作智能体间多样化且自适应的策略生成。
原文摘要 · Abstract (English)
In recent years, diversity has emerged as a useful mechanism to enhance the efficiency of multi-agent reinforcement learning (MARL). However, existing methods predominantly focus on designing policies based on individual agent characteristics, often neglecting the interplay and mutual influence among agents during policy formation. To address this gap, we propose Competitive Diversity through Constructive Conflict (CoDiCon), a novel approach that incorporates competitive incentives into cooperative scenarios to encourage policy exchange and foster strategic diversity among agents. Drawing inspiration from sociological research, which highlights the benefits of moderate competition and constructive conflict in group decision-making, we design an intrinsic reward mechanism using ranking features to introduce competitive motivations. A centralized intrinsic reward module generates and distributes varying reward values to agents, ensuring an effective balance between competition and cooperation. By optimizing the parameterized centralized reward module to maximize environmental rewards, we reformulate the constrained bilevel optimization problem to align with the original task objectives. We evaluate our algorithm against state-of-the-art methods in the SMAC and GRF environments. Experimental results demonstrate that CoDiCon achieves superior performance, with competitive intrinsic rewards effectively promoting diverse and adaptive strategies among cooperative agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。