用世界模型模拟轨迹,高效生成多样协作伙伴
Efficient Generation of Diverse Cooperative Agents with World Models
- 用学习的世界模型生成环境动态轨迹,替代真实采样
- 仅需单次轨迹即可完成训练,性能媲美原有方法
- 适合需要大量协作伙伴的零样本协同场景
零样本协同(ZSC)智能体训练中的主要瓶颈在于生成具有多样化协作方式的伙伴智能体。现有交叉对弈最小化(XPM)方法在群体生成中计算成本高、采样效率低,因训练目标需采样多种轨迹,且每个伙伴均从头训练,尽管它们学习的是同一协调任务。本文提出利用环境动力学模型生成的模拟轨迹,显著加速XPM训练过程。我们引入XPM-WM框架,通过学习的世界模型(WM)生成用于XPM的模拟轨迹。结果表明,使用模拟轨迹可免除多次轨迹采样需求,且所生成伙伴在多样性与协作性能上达到先前方法水平,同时大幅提升样本效率并支持更大规模伙伴群体生成。
原文摘要 · Abstract (English)
A major bottleneck in the training process for Zero-Shot Coordination (ZSC) agents is the generation of partner agents that are diverse in collaborative conventions. Current Cross-play Minimization (XPM) methods for population generation can be very computationally expensive and sample inefficient as the training objective requires sampling multiple types of trajectories. Each partner agent in the population is also trained from scratch, despite all of the partners in the population learning policies of the same coordination task. In this work, we propose that simulated trajectories from the dynamics model of an environment can drastically speed up the training process for XPM methods. We introduce XPM-WM, a framework for generating simulated trajectories for XPM via a learned World Model (WM). We show XPM with simulated trajectories removes the need to sample multiple trajectories. In addition, we show our proposed method can effectively generate partners with diverse conventions that match the performance of previous methods in terms of SP population training reward as well as training partners for ZSC agents. Our method is thus, significantly more sample efficient and scalable to a larger number of partners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。