arXiv:2603.07253cs.MAcs.AI2026-03

让智能体学会在目标不同时判断是否合作,提升开放环境协作效率。

Learning When to Cooperate Under Heterogeneous Goals

  • 用模仿学习与强化学习分层结合,自动决策何时合作或独行。
  • 在两个协作环境中表现优于基线方法,尤其在目标重叠度低时优势明显。
  • 能预测队友行为的辅助模块,信息越少越关键,适配复杂多变场景。

人类合作智慧的核心在于识别有利合作时机,并判断任务是否更适合独自完成。现有机器灵活协作研究尚未充分探索这一元层面问题,而这对异构开放环境中的成功协作至关重要。本文将典型的即兴团队(Ad Hoc Teamwork, AHT)设定拓展为包含异质目标的场景,其中目标可能重叠也可能不重叠。提出一种基于模仿学习与强化学习分层结合的新策略学习方法,在两个扩展的协作环境中均优于基线方法。此外,研究发现一个辅助模块——通过预测队友行为来建模同伴——对性能的影响与可观察到的队友目标信息量呈反比,说明在信息不足时该模块尤为重要。

原文摘要 · Abstract (English)

A significant element of human cooperative intelligence lies in our ability to identify opportunities for fruitful collaboration; and conversely to recognise when the task at hand is better pursued alone. Research on flexible cooperation in machines has left this meta-level problem largely unexplored, despite its importance for successful collaboration in heterogeneous open-ended environments. Here, we extend the typical Ad Hoc Teamwork (AHT) setting to incorporate the idea of agents having heterogeneous goals that in any given scenario may or may not overlap. We introduce a novel approach to learning policies in this setting, based on a hierarchical combination of imitation and reinforcement learning, and show that it outperforms baseline methods across extended versions of two cooperative environments. We also investigate the contribution of an auxiliary component that learns to model teammates by predicting their actions, finding that its effect on performance is inversely related to the amount of observable information about teammate goals.

协作智能强化学习异质目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。