arXiv:2506.19417cs.LGcs.MA2025-06被引 2

让多智能体在稀疏奖励下聚焦未探索区域,提升协作效率

Focusing Influence Mechanism for Multi-Agent Reinforcement Learning

  • 基于熵准则引导智能体聚焦未探索状态空间
  • 用资格迹实现多智能体持续一致地影响同一区域
  • 适用于稀疏奖励下的协作强化学习任务

在稀疏奖励下,协作式多智能体强化学习(MARL)仍面临根本性挑战,因智能体往往无法集中其影响,导致探索协调不足。为此,我们提出聚焦影响机制(Focusing Influence Mechanism, FIM),通过熵基准则促使智能体关注状态空间中未充分探索的部分,同时利用资格迹使多个智能体在有益时能持续一致地对同一区域施加影响,从而促进协同且持久的联合行为。通过强调状态空间中的未探索区域,FIM即使在极端稀疏奖励下也能实现更高效、结构化的探索。在多种MARL基准测试中,FIM始终优于强基线方法。

原文摘要 · Abstract (English)

Cooperative multi-agent reinforcement learning (MARL) under sparse rewards remains fundamentally challenging because agents often fail to concentrate their influence, leading to insufficiently coordinated exploration. To address this, we propose the Focusing Influence Mechanism (FIM), a framework that encourages agents to focus their influence on under-explored parts of the state space through an entropy-based criterion, while leveraging eligibility traces to enable multiple agents to consistently align and sustain their influence on the same parts of the state space when beneficial, thereby promoting coordinated and persistent joint behavior. By emphasizing under-explored regions of the state space, FIM facilitates more efficient and structured exploration even under extremely sparse rewards. Across diverse MARL benchmarks, FIM consistently improves cooperative performance over strong baselines.

多智能体强化学习协作探索稀疏奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。