arXiv:2508.06336cs.LGcs.AI2025-08被引 3

无需预设伙伴,自动生成适应性队友,提升协作鲁棒性。

Unsupervised Partner Design Enables Robust Ad-hoc Teamwork

  • 动态生成训练伙伴,按可学习性自动选择,无需预训练群体。
  • 在多个任务中表现优于有无群体的基线方法,人机测试中得分更高。
  • 适合需要灵活协作的多智能体系统,如游戏或机器人团队。

我们提出无监督伙伴设计(UPD),一种无需预设群体的多智能体强化学习方法,用于实现稳健的临时协作。UPD 实时生成训练伙伴,并基于可学习性准则自适应选择,避免了对预训练伙伴群体或手动调参的需求。该简单机制有效实现了伙伴多样性,并可在具备过程生成器时扩展为联合伙伴-环境选择。在基于关卡的觅食、Overcooked-AI 及 Overcooked 泛化挑战任务中,UPD 持续优于各类基于群体与无群体的基线方法。在人类-AI 用户研究中,经 UPD 训练的智能体获得更高回报,且被评价为更具适应性、更像人类、更不易令人沮丧,优于所有对比方法。

原文摘要 · Abstract (English)

We introduce Unsupervised Partner Design (UPD), a population-free multi-agent reinforcement learning method for robust ad-hoc teamwork. UPD generates training partners on-the-fly and selects them adaptively based on a learnability criterion, removing the need for pre-trained partner populations or manual parameter tuning. We show that this simple mechanism enables effective partner diversity and can be extended to joint partner-environment selection when a procedural level generator is available. Across Level-Based Foraging, Overcooked-AI, and the Overcooked Generalisation Challenge, UPD consistently achieves strong performance compared to both population-based and population-free baselines. In a human-AI user study, agents trained with UPD achieve higher returns and are rated as more adaptive, more human-like, and less frustrating than all evaluated baseline methods.

多智能体强化学习协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。