让智能体零样本协作,用通用策略提升适应新队友能力
Zero-Shot Coordination in Ad Hoc Teams with Generalized Policy Improvement and Difference Rewards
- 整合所有预训练策略,通过泛化策略改进实现跨团队协作
- 在三个仿真环境和真实机器人场景中均实现零样本成功协作
- 适合需要快速适配陌生队友的多智能体系统应用
现实中的多智能体系统可能需要在未知队友的情况下进行即兴协作,以零样本方式完成任务。以往方法通常基于对新队友的建模选择预训练策略,或预训练单一鲁棒策略。本文提出在零样本迁移设置下充分利用所有预训练策略。我们将问题形式化为即兴多智能体马尔可夫决策过程,并提出两种关键机制:泛化策略改进与差异奖励,以实现不同团队间的高效知识迁移。实验表明,所提算法GPAT在三个仿真环境(合作觅食、捕食者-猎物、Overcooked)中成功实现对新团队的零样本迁移,并在真实多机器人场景中验证有效性。
原文摘要 · Abstract (English)
Real-world multi-agent systems may require ad hoc teaming, where an agent must coordinate with other previously unseen teammates to solve a task in a zero-shot manner. Prior work often either selects a pretrained policy based on an inferred model of the new teammates or pretrains a single policy that is robust to potential teammates. Instead, we propose to leverage all pretrained policies in a zero-shot transfer setting. We formalize this problem as an ad hoc multi-agent Markov decision process and present a solution that uses two key ideas, generalized policy improvement and difference rewards, for efficient and effective knowledge transfer between different teams. We empirically demonstrate that our algorithm, Generalized Policy improvement for Ad hoc Teaming (GPAT), successfully enables zero-shot transfer to new teams in three simulated environments: cooperative foraging, predator-prey, and Overcooked. We also demonstrate our algorithm in a real-world multi-robot setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。