让智能体学会与不同奖励设计的伙伴协作,提升稀疏奖励下的合作效率。
Zero Shot Coordination for Sparse Reward Tasks with Diverse Reward Shapings

- 用随机奖励设计训练多方法集成,增强对未知伙伴的适应能力。
- 在Overcooked中相比基线算法提升62.2%至119.2%的稀疏奖励表现。
- 适合研究零样本协作、多智能体强化学习的开发者和研究人员。
许多多智能体强化学习(MARL)智能体无法有效适应与使用相同目标但训练种子、算法等不同的智能体协作。这即为零样本协作(ZSC)问题,旨在训练智能体以与未知智能体良好协作。尽管已有研究在表格场景和简单游戏如汉诺比中取得优异成果,但现有方案仅考虑训练智能体与未来伙伴具有完全相同的奖励。这在现实中不成立——智能体往往面临相同稀疏目标但奖励设计不同的协作对象。为此,本文提出通过四种选择算法随机选取奖励设计,训练一组方法的集成。在Overcooked环境中的实验表明,与具有相同稀疏奖励但不同奖励设计的智能体协作时,性能相比基线ZSC算法提升62.2%至119.2%。
原文摘要 · Abstract (English)
Many Multi-Agent Reinforcement Learning (MARL) agents fail to adapt properly to cooperating with agents trained with the same objectives but different seeds, algorithms, or other training differences. This is the problem of Zero-Shot Coordination (ZSC), which focuses on training agents to cooperate well with unknown agents. ZSC has been studied for a variety of tabular cases and simple games such as Hanabi, achieving excellent results. However, existing solutions to ZSC only consider identical rewards for your trained agents and all future partners. This is not realistic for the trained agents, as they do not consider the problem of cooperating with agents that have identical sparse objectives but shape the rewards for those objectives in different manner. To address this issue, we show how to train an ensemble of methods using randomized reward shapings chosen using 4 selection algorithms. Experiments done on the Overcooked environment demonstrate consistent improvements of 62.2%-119.2% in sparse reward over baseline ZSC algorithms when playing with agents that have identical sparse rewards but different reward shapings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。