用扩散模型让智能体预判并适应临时队友,实现多样协作。
PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork
- 基于扩散模型捕捉多模式合作行为
- 在三个环境中显著超越现有方法
- 适合需要动态协作的复杂场景
临时团队协作(AHT)要求智能体与未曾见过的队友协同工作,这对许多现实应用至关重要。核心挑战在于训练一个能实时预测并适应未知队友的主体智能体。传统基于强化学习的方法优化单一期望回报,常导致策略坍缩为单一主导行为,无法捕捉AHT中固有的多模态协作模式。本文提出PADiff,一种基于扩散模型的方法,能够捕获智能体的多模态行为,解锁其与队友的多样化协作方式。然而,标准扩散模型缺乏在高度非平稳的AHT场景中进行预测与适应的能力。为此,我们提出一种新型扩散策略,将关于队友的关键预测信息融入去噪过程。在三个协作环境中的大量实验表明,PADiff显著优于现有AHT方法。
原文摘要 · Abstract (English)
Ad hoc teamwork (AHT) requires agents to collaborate with previously unseen teammates, which is crucial for many real-world applications. The core challenge of AHT is to develop an ego agent that can predict and adapt to unknown teammates on the fly. Conventional RL-based approaches optimize a single expected return, which often causes policies to collapse into a single dominant behavior, thus failing to capture the multimodal cooperation patterns inherent in AHT. In this work, we introduce PADiff, a diffusion-based approach that captures agent's multimodal behaviors, unlocking its diverse cooperation modes with teammates. However, standard diffusion models lack the ability to predict and adapt in highly non-stationary AHT scenarios. To address this limitation, we propose a novel diffusion-based policy that integrates critical predictive information about teammates into the denoising process. Extensive experiments across three cooperation environments demonstrate that PADiff outperforms existing AHT methods significantly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。