arXiv:2511.12754cs.AIcs.LG2025-11NeurIPS被引 4

让智能体实时学习并适应新伙伴的策略,提升协作效率。

Adaptively Coordinating with Novel Partners via Learned Latent Strategies

  • 用变分自编码器构建潜在策略空间,聚类识别不同策略类型。
  • 在Overcooked环境中,对新人类和智能体伙伴表现优于现有方法。
  • 在线适应新伙伴时,动态调整策略估计,适合复杂协作场景。

适应是异构团队有效协作的核心。在人机协作中,智能体需实时适应人类伙伴的独特偏好与策略变化,尤其在时间紧迫、策略空间复杂的任务中更具挑战性。本文提出一种策略条件化的合作者框架,通过变分自编码器从智能体轨迹数据中学习潜在策略空间,利用聚类识别不同策略类型,并训练合作者在每类策略下生成对应响应。针对在线适应新伙伴,采用固定共享后悔最小化算法,在交互中动态推断并调整对伙伴策略的估计。在改进版Overcooked协作烹饪环境中的实验及在线用户研究显示,该方法在与新的人类或智能体伙伴配对时,性能达到当前最优水平。

原文摘要 · Abstract (English)

Adaptation is the cornerstone of effective collaboration among heterogeneous team members. In human-agent teams, artificial agents need to adapt to their human partners in real time, as individuals often have unique preferences and policies that may change dynamically throughout interactions. This becomes particularly challenging in tasks with time pressure and complex strategic spaces, where identifying partner behaviors and selecting suitable responses is difficult. In this work, we introduce a strategy-conditioned cooperator framework that learns to represent, categorize, and adapt to a broad range of potential partner strategies in real-time. Our approach encodes strategies with a variational autoencoder to learn a latent strategy space from agent trajectory data, identifies distinct strategy types through clustering, and trains a cooperator agent conditioned on these clusters by generating partners of each strategy type. For online adaptation to novel partners, we leverage a fixed-share regret minimization algorithm that dynamically infers and adjusts the partner's strategy estimation during interaction. We evaluate our method in a modified version of the Overcooked domain, a complex collaborative cooking environment that requires effective coordination among two players with a diverse potential strategy space. Through these experiments and an online user study, we demonstrate that our proposed agent achieves state of the art performance compared to existing baselines when paired with novel human, and agent teammates.

协作智能体策略适应人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。