自动设计奖励信号,让多智能体在稀疏奖励下更稳定高效学习。
ARMS: Automatic Reward Shaping for Sparse-Reward Multi-Agent Reinforcement Learning

- 基于轨迹排序自动生成密集奖励信号,共享参数提升效率。
- 在复杂路径寻路任务中,奖励越稀疏、智能体越多,效果越显著。
- 首次从博弈论角度保证均衡结构不变,适合研究多智能体协作的学者。
稀疏奖励是多智能体强化学习(MARL)的主要瓶颈,同时学习导致非平稳性,使奖励设计尤为困难。奖励塑造可加速学习,但在多智能体场景中必须保持问题的战略结构,而非仅优化短期收益。本文提出自动多智能体奖励塑造框架ARMS,通过轨迹排序从稀疏环境奖励中学习密集塑造信号。由于单智能体轨迹排序不直接适用于MARL,我们通过条件最优响应推理重新定义策略不变性,并证明:在特定条件下,使用塑造奖励可保持每个智能体在对手策略固定时的最优响应集,从而保全纳什均衡集。基于此,ARMS在策略学习与奖励学习间交替进行,并共享塑造参数以提高效率。在部分可观测多智能体路径寻路环境中,ARMS在奖励稀疏性和智能体数量增加时均提升采样效率,具备对未见环境的泛化能力,并揭示一种多智能体特有失败模式:探索受限与耦合的策略-奖励动态引发振荡行为。增强探索可缓解该现象并稳定学习。据我们所知,ARMS是首个基于博弈论均衡保持结果设计的自动奖励塑造框架。
原文摘要 · Abstract (English)
Sparse rewards are a major bottleneck in multi-agent reinforcement learning (MARL), where simultaneous learning induces non-stationarity and makes reward design especially delicate. Reward shaping can accelerate learning, but in the multi-agent setting it must preserve the strategic structure of the problem rather than merely improve short-term optimization. We propose Automatic Reward-shaping in Multi-agent Systems (ARMS), a self-supervised reward shaping framework for MARL that learns dense shaping signals from sparse environmental rewards through trajectory ranking. Since single-agent trajectory-ranking guarantees do not directly transfer to MARL, we reformulate policy invariance through conditional best-response reasoning, and show that if certain conditions hold, then using shaping rewards preserves each agent's best-response set under fixed opponent policies, and consequently preserve the set of Nash equilibria. Guided by this perspective, ARMS alternates between policy learning and reward learning while sharing shaping parameters across agents for efficiency. Experiments in a partially observable multi-agent pathfinding domain show that ARMS improves sampling efficiency under increasing reward sparsity and agent count, generalizes to unseen environments, and reveals a MARL-specific failure mode in which limited exploration and coupled policy--reward dynamics induce oscillatory behavior. Increasing exploration mitigates this effect and stabilizes learning. To the best of our knowledge, ARMS is the first automatic reward shaping framework for MARL whose design is motivated by a game-theoretic equilibrium-preservation result.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。