arXiv:2505.08630cs.LG2025-05

通过影响范围计算,精准分配多智能体奖励并提升探索效率

Credit Assignment and Efficient Exploration based on Influence Scope in Multi-agent Reinforcement Learning

  • 基于状态属性的影响范围计算各智能体作用域
  • 在稀疏奖励环境下显著优于现有基线方法
  • 适合需要高效协作与奖励分配的多智能体场景

在稀疏奖励场景中训练协作智能体对多智能体强化学习(MARL)构成重大挑战。由于每一步缺乏明确的动作反馈,以往方法在智能体间难以实现精确的信用分配和有效探索。本文提出一种新方法,通过计算单个智能体对状态特定维度/属性的影响范围(ISA),利用智能体动作与状态属性间的相互依赖关系,实现信用分配并限定每个智能体的探索空间。我们在多种稀疏奖励多智能体场景中评估了ISA,结果表明该方法显著优于当前最优基线。

原文摘要 · Abstract (English)

Training cooperative agents in sparse-reward scenarios poses significant challenges for multi-agent reinforcement learning (MARL). Without clear feedback on actions at each step in sparse-reward setting, previous methods struggle with precise credit assignment among agents and effective exploration. In this paper, we introduce a novel method to deal with both credit assignment and exploration problems in reward-sparse domains. Accordingly, we propose an algorithm that calculates the Influence Scope of Agents (ISA) on states by taking specific value of the dimensions/attributes of states that can be influenced by individual agents. The mutual dependence between agents' actions and state attributes are then used to calculate the credit assignment and to delimit the exploration space for each individual agent. We then evaluate ISA in a variety of sparse-reward multi-agent scenarios. The results show that our method significantly outperforms the state-of-art baselines.

多智能体信用分配强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。