arXiv:2511.00034cs.MAcs.LG2025-11被引 2

decentralized reward shaping难在协调,无法提升多智能体合作性能。

On the Fundamental Limitations of Decentralized Learnable Reward Shaping in Cooperative Multi-Agent Reinforcement Learning

  • 各智能体独立学习奖励函数,完全去中心化
  • 平均奖励仅-24.20,比中心化方法低26.12分
  • 局部效率高但全局协作差,揭示协调悖论

近期可学习奖励塑造在单智能体强化学习中展现潜力,但在合作式多智能体场景下的去中心化方法效果仍不明确。本文提出完全去中心化的DMARL-RSA系统,每个智能体独立学习奖励塑造,并在simple_spread_v3环境中评估。尽管采用先进学习机制,其平均奖励仅为-24.20 ± 0.09,远低于中心化训练的MAPPO(1.92 ± 0.87),差距达26.12分。其表现与独立学习(IPPO: -23.19 ± 0.96)相当,表明复杂奖励塑造无法克服去中心化协作的根本限制。有趣的是,去中心化方法在地标覆盖上表现更优(DMARL-RSA: 0.888 ± 0.029,IPPO: 0.960 ± 0.045,共3个地标),但整体性能低于中心化方法(0.273 ± 0.008),暴露局部优化与全局目标之间的协调悖论。分析识别出三大障碍:(1) 策略更新并发导致非平稳性,(2) 信用分配复杂度呈指数增长,(3) 个体奖励优化与全局目标错位。这些结果确立了去中心化奖励学习的实证极限,强调有效多智能体协作需依赖中心化协调。

原文摘要 · Abstract (English)

Recent advances in learnable reward shaping have shown promise in single-agent reinforcement learning by automatically discovering effective feedback signals. However, the effectiveness of decentralized learnable reward shaping in cooperative multi-agent settings remains poorly understood. We propose DMARL-RSA, a fully decentralized system where each agent learns individual reward shaping, and evaluate it on cooperative navigation tasks in the simple_spread_v3 environment. Despite sophisticated reward learning, DMARL-RSA achieves only -24.20 +/- 0.09 average reward, compared to MAPPO with centralized training at 1.92 +/- 0.87 -- a 26.12-point gap. DMARL-RSA performs similarly to simple independent learning (IPPO: -23.19 +/- 0.96), indicating that advanced reward shaping cannot overcome fundamental decentralized coordination limitations. Interestingly, decentralized methods achieve higher landmark coverage (0.888 +/- 0.029 for DMARL-RSA, 0.960 +/- 0.045 for IPPO out of 3 total) but worse overall performance than centralized MAPPO (0.273 +/- 0.008 landmark coverage) -- revealing a coordination paradox between local optimization and global performance. Analysis identifies three critical barriers: (1) non-stationarity from concurrent policy updates, (2) exponential credit assignment complexity, and (3) misalignment between individual reward optimization and global objectives. These results establish empirical limits for decentralized reward learning and underscore the necessity of centralized coordination for effective multi-agent cooperation.

多智能体强化学习奖励塑造去中心化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。