MAGIC通过多步因果影响提升多智能体协作效率
MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning

- 用反事实干预估计智能体间多步动作影响
- 平均性能比现有方法提升26.9%和10.1%
- 适合需要强协作的复杂任务场景
多智能体强化学习(MARL)的核心挑战在于设计能有效促进智能体间协作的学习信号。这要求估算一个智能体当前动作对未来交互步骤中队友的影响。为此,我们提出多步优势门控干预因果MARL(MAGIC),通过反事实动作干预比较真实与假设情境下队友未来状态的差异,进而估算多步动作影响,并利用优势值作为门控机制,引导探索聚焦于符合任务目标的有益行为。在多智能体粒子环境(MPE)、星际争霸微操基准(SMAC与SMACv2)上的实验表明,MAGIC consistently优于现有领先方法,在两个基准上分别实现平均相对性能提升26.9%和10.1%。
原文摘要 · Abstract (English)
A key challenge in multi-agent reinforcement learning (MARL) lies in designing learning signals that effectively promote coordination among agents. Designing such signals requires estimating how one agent's current action affects its teammates over future interaction steps. To address this, we introduce Multi-step Advantage-Gated Interventional Causal MARL (MAGIC), a framework that estimates multi-step action effects between agents and selectively converts them into intrinsic rewards. MAGIC uses counterfactual action interventions to compare teammate futures under factual and counterfactual branches, and introduces a gate based on advantage to direct exploration toward beneficial behaviors aligned with the task goal. Experiments on Multi-Agent Particle Environments (MPE) and StarCraft micromanagement benchmarks (SMAC and SMACv2) show that MAGIC consistently outperforms leading prior methods, with average relative final performance improvements of 26.9% and 10.1%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。