arXiv:2504.12714cs.MAcs.AI2025-04ICML被引 19

让智能体在多种环境间协作,实现与陌生伙伴零样本协同。

Cross-environment Cooperation Enables Zero-shot Multi-agent Coordination

  • 通过跨环境协作训练,让智能体学会通用合作技能。
  • 在真实人类协作中表现优于基线模型,支持零样本适应新伙伴。
  • 无需人类数据,适合开发可与人自然互动的通用协作智能体。

零样本协调(ZSC)是人机兼容人工智能的关键能力,指智能体在未见过的新伙伴面前仍能有效协作。现有研究多聚焦于单一任务的协作训练,但此类模型难以泛化到相似新任务。本文提出跨环境协作(CEC)范式,通过在包含数十亿个可解协作挑战的多样化环境中训练,使智能体学习通用合作策略。我们构建了两个基于Jax的程序化生成器,用于生成海量协调难题。实验表明,采用该方法的智能体在与真实人类协作时,不仅定量表现更优,且定性上更具适应性。结果说明,在多种独特场景中协作能促使智能体发展出通用协作规范,从而有效应对不同伙伴与任务。这为设计无需人类数据即可与人类自然交互的通用协作智能体提供了新路径。

原文摘要 · Abstract (English)

Zero-shot coordination (ZSC), the ability to adapt to a new partner in a cooperative task, is a critical component of human-compatible AI. While prior work has focused on training agents to cooperate on a single task, these specialized models do not generalize to new tasks, even if they are highly similar. Here, we study how reinforcement learning on a distribution of environments with a single partner enables learning general cooperative skills that support ZSC with many new partners on many new problems. We introduce two Jax-based, procedural generators that create billions of solvable coordination challenges. We develop a new paradigm called Cross-Environment Cooperation (CEC), and show that it outperforms competitive baselines quantitatively and qualitatively when collaborating with real people. Our findings suggest that learning to collaborate across many unique scenarios encourages agents to develop general norms, which prove effective for collaboration with different partners. Together, our results suggest a new route toward designing generalist cooperative agents capable of interacting with humans without requiring human data.

多智能体零样本协作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。