arXiv:2511.17165cs.AIcs.LG2025-11

提出MIR机制,让多智能体在稀疏奖励下主动探索影响队友的动作。

MIR: Efficient Exploration in Episodic Multi-Agent Reinforcement Learning via Mutual Intrinsic Reward

  • 设计互斥内在奖励,鼓励智能体探索影响队友的动作
  • 在稀疏奖励环境下,团队性能显著优于现有方法
  • 适用于需要协作探索的多智能体任务

回合制奖励给强化学习带来重大挑战。尽管内在奖励方法在单智能体场景中表现良好,但在多智能体强化学习(MARL)中仍面临困难。主要源于两点:(1) 联合动作轨迹呈指数级稀疏,导致奖励难以获取;(2) 现有方法常忽略能影响团队状态的联合动作。为此,本文提出互斥内在奖励(MIR),一种针对稀疏奖励如回合制奖励的高效增强策略。MIR激励个体智能体探索影响队友的动作,与原有策略结合后有效激发团队探索,提升算法性能。为全面验证,我们将代表性单智能体环境MiniGrid扩展为MiniGrid-MA,构建一系列具有稀疏奖励的MARL环境。评估结果表明,所提方法在MiniGrid-MA设置下优于当前最先进方法。

原文摘要 · Abstract (English)

Episodic rewards present a significant challenge in reinforcement learning. While intrinsic reward methods have demonstrated effectiveness in single-agent rein-forcement learning scenarios, their application to multi-agent reinforcement learn-ing (MARL) remains problematic. The primary difficulties stem from two fac-tors: (1) the exponential sparsity of joint action trajectories that lead to rewards as the exploration space expands, and (2) existing methods often fail to account for joint actions that can influence team states. To address these challenges, this paper introduces Mutual Intrinsic Reward (MIR), a simple yet effective enhancement strategy for MARL with extremely sparse rewards like episodic rewards. MIR incentivizes individual agents to explore actions that affect their teammates, and when combined with original strategies, effectively stimulates team exploration and improves algorithm performance. For comprehensive experimental valida-tion, we extend the representative single-agent MiniGrid environment to create MiniGrid-MA, a series of MARL environments with sparse rewards. Our evalu-ation compares the proposed method against state-of-the-art approaches in the MiniGrid-MA setting, with experimental results demonstrating superior perfor-mance.

多智能体强化学习稀疏奖励探索机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。