arXiv:2412.14779cs.MAcs.AI2024-12被引 1

解决多智能体稀疏奖励下的学习难题,让每个智能体在合适时间点获得合理奖励反馈。

Agent-Temporal Credit Assignment for Optimal Policy Preservation in Sparse Multi-Agent Reinforcement Learning

  • 将全局稀疏奖励分解为时间步与个体贡献,实现精准的信用分配。
  • 实验表明学习更稳定且加速,性能媲美甚至超越传统多智能体方法。
  • 理论保证最优策略不变,适合长周期、延迟奖励的复杂协作任务。

在多智能体环境中,由于全局奖励稀疏或延迟,智能体难以学习到最优策略,尤其是在长时序任务中,难以评估中间步骤的行为价值。我们提出一种名为时序-智能体奖励重分配(TAR²)的新方法,旨在解决代理-时序信用分配问题,通过在时间和智能体间重新分配稀疏奖励。TAR²将稀疏的全局奖励分解为特定时间步的奖励,并计算各智能体对这些奖励的贡献。理论上证明了TAR²等价于基于势能的奖励塑形,确保最优策略不变。实验结果表明,TAR²可稳定并加速学习过程。此外,当TAR²与单智能体强化学习算法结合时,其表现可媲美甚至优于传统多智能体强化学习方法。

原文摘要 · Abstract (English)

In multi-agent environments, agents often struggle to learn optimal policies due to sparse or delayed global rewards, particularly in long-horizon tasks where it is challenging to evaluate actions at intermediate time steps. We introduce Temporal-Agent Reward Redistribution (TAR$^2$), a novel approach designed to address the agent-temporal credit assignment problem by redistributing sparse rewards both temporally and across agents. TAR$^2$ decomposes sparse global rewards into time-step-specific rewards and calculates agent-specific contributions to these rewards. We theoretically prove that TAR$^2$ is equivalent to potential-based reward shaping, ensuring that the optimal policy remains unchanged. Empirical results demonstrate that TAR$^2$ stabilizes and accelerates the learning process. Additionally, we show that when TAR$^2$ is integrated with single-agent reinforcement learning algorithms, it performs as well as or better than traditional multi-agent reinforcement learning methods.

多智能体强化学习信用分配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。