通过同伴激励提升多智能体探索效率
PIMAEX: Multi-Agent Exploration through Peer Incentivization
- 设计同伴激励奖励机制,促进智能体相互影响以发现新状态
- 在存在误导奖励的环境中显著优于基线方法
- 适合研究多智能体协作与探索难题的学者
尽管单智能体强化学习中的探索问题已得到广泛研究,但多智能体强化学习中的探索仍缺乏关注。本文提出一种受内在好奇心和影响力奖励启发的同伴激励奖励函数——PIMAEX(Peer-Incentivized Multi-Agent Exploration),旨在通过鼓励智能体相互施加影响,提高遇到新状态的概率,从而改善多智能体环境下的探索性能。进一步结合PIMAEX-Communication训练算法,该算法引入通信通道使智能体能相互影响。在专为挑战探索与利用权衡及信用分配问题而设计的局部可观测环境Consume/Explore中进行评估,结果表明,使用PIMAEX奖励与通信机制的智能体在探索效率上显著优于未使用该机制的基线方法。
原文摘要 · Abstract (English)
While exploration in single-agent reinforcement learning has been studied extensively in recent years, considerably less work has focused on its counterpart in multi-agent reinforcement learning. To address this issue, this work proposes a peer-incentivized reward function inspired by previous research on intrinsic curiosity and influence-based rewards. The \textit{PIMAEX} reward, short for Peer-Incentivized Multi-Agent Exploration, aims to improve exploration in the multi-agent setting by encouraging agents to exert influence over each other to increase the likelihood of encountering novel states. We evaluate the \textit{PIMAEX} reward in conjunction with \textit{PIMAEX-Communication}, a multi-agent training algorithm that employs a communication channel for agents to influence one another. The evaluation is conducted in the \textit{Consume/Explore} environment, a partially observable environment with deceptive rewards, specifically designed to challenge the exploration vs.\ exploitation dilemma and the credit-assignment problem. The results empirically demonstrate that agents using the \textit{PIMAEX} reward with \textit{PIMAEX-Communication} outperform those that do not.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。