arXiv:2601.21523cs.LG2026-01

通过局部奖励与依赖图实现多智能体强化学习中的精准责任分配

Explicit Credit Assignment through Local Rewards and Dependence Graphs in Multi-Agent Reinforcement Learning

  • 构建智能体间交互图,精细区分各智能体贡献
  • 实验显示该方法在合作任务中优于传统局部与全局奖励
  • 适合需要高效协作的多智能体系统设计

为促进多智能体强化学习中的合作,可将所有智能体的奖励聚合形成全局奖励,即完全合作设置。然而,全局奖励通常噪声较大,因包含所有智能体的贡献,需在信用分配过程中加以分辨。相反,使用局部奖励虽能加快学习速度,因分离了智能体贡献,但可能导致次优结果,因智能体仅短视地优化自身奖励而忽略全局最优。本文提出一种结合两者优势的方法:利用智能体间的交互图,以比全局奖励更细粒度的方式识别个体贡献,同时缓解局部奖励带来的合作问题。我们还提出一种实用的近似该图的方法。实验表明该方法具有灵活性,能在传统局部与全局奖励设置上取得改进。

原文摘要 · Abstract (English)

To promote cooperation in Multi-Agent Reinforcement Learning, the reward signals of all agents can be aggregated together, forming global rewards that are commonly known as the fully cooperative setting. However, global rewards are usually noisy because they contain the contributions of all agents, which have to be resolved in the credit assignment process. On the other hand, using local reward benefits from faster learning due to the separation of agents' contributions, but can be suboptimal as agents myopically optimize their own reward while disregarding the global optimality. In this work, we propose a method that combines the merits of both approaches. By using a graph of interaction between agents, our method discerns the individual agent contribution in a more fine-grained manner than a global reward, while alleviating the cooperation problem with agents' local reward. We also introduce a practical approach for approximating such a graph. Our experiments demonstrate the flexibility of the approach, enabling improvements over the traditional local and global reward settings.

多智能体强化学习信用分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。