让AI通过观察快速适应陌生伙伴,无需参数更新。
CooT: Learning to Coordinate In-Context with Coordination Transformers
- 利用上下文学习观察伙伴行为,实时调整自身动作。
- 在厨艺和足球任务中表现优于现有方法,稳定快速适应。
- 适合需要即时协作的现实场景,如人机协同工作。
在多智能体系统中,与陌生伙伴高效协作仍是重大挑战。现有方法如群体训练虽提升鲁棒性,但难以在训练分布外高效适应;微调因交互成本高,在少样本场景下不实用。为此,我们提出CooT框架,利用上下文学习(ICL)实现实时伙伴适应。不同于以往聚焦任务泛化的ICL,CooT专为跨多样化伙伴行为泛化设计。在偏好行为轨迹上训练后,仅通过观察即可对齐行动与伙伴意图。在Overcooked和Google Research Football两个复杂多智能体基准上评估显示,CooT始终优于群体方法、基于梯度微调及Meta-RL基线,在无参数更新条件下实现稳定且快速适应。人类评估也表明其为更受欢迎的合作对象,消融实验确认其可快速适应新伙伴并在伙伴突变时保持稳定,适用于真实世界人机协作。
原文摘要 · Abstract (English)
Effective coordination among unfamiliar partners remains a major challenge in multi-agent systems. Existing approaches, such as population-based methods, improve robustness through diversity but often lack mechanisms for efficient adaptation beyond training distribution. Moreover, fine-tuning is impractical in few-shot settings due to its high interaction cost. To address these limitations, we propose CooT, a framework that leverages in-context learning (ICL) for real-time partner adaptation. Unlike prior ICL approaches that focus on task generalization, CooT is designed to generalize across diverse partner behaviors. Trained on trajectories from behavior-preferring agents, it learns to align actions with partner intentions purely through observation. We evaluate CooT on two challenging multi-agent benchmarks: Overcooked and Google Research Football. Results show that CooT consistently outperforms population-based methods, gradient-based fine-tuning, and Meta-RL baselines, achieving stable and rapid adaptation without parameter updates. Human evaluations also identify CooT as a preferred collaborator, and our ablations confirm its ability to adapt quickly to new partners and remain stable under sudden partner changes, making it reliable for real-world human-AI collaboration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。