用检索增强的LLM建模临时队友行为,提升协作适应能力。
ReCollab: Retrieval-Augmented LLMs for Cooperative Ad-hoc Teammate Modeling
- 基于轨迹特征构建行为评分卡,用LLM分类队友类型。
- 引入检索增强生成,使模型在不同场景下更稳定地适应队友。
- 适合研究多智能体协作与动态伙伴建模的学者参考。
临时团队协作(AHT)要求智能体推断未见过的队友行为并相应调整自身策略。传统方法依赖固定概率模型或分类器,在部分可观测和交互有限的情况下易失效。大语言模型(LLMs)提供灵活替代方案:通过将短时行为轨迹映射为高层假设,可作为队友行为的世界模型。我们提出 extsc{Collab},一种基于语言的框架,利用轨迹特征导出的行为评分卡对伙伴类型进行分类;进一步扩展为 extsc{ReCollab},引入检索增强生成(RAG),通过实例化轨迹稳定推理过程。在协作式过厨房(Overcooked)环境中, extsc{Collab} 能有效区分队友类型,而 extsc{ReCollab} 在多种布局下持续提升适应性,实现分类准确率与回合回报之间的帕累托最优权衡。结果表明,LLMs 可作为 AHT 中的行为世界模型,且检索增强在复杂协调场景中至关重要。
原文摘要 · Abstract (English)
Ad-hoc teamwork (AHT) requires agents to infer the behavior of previously unseen teammates and adapt their policy accordingly. Conventional approaches often rely on fixed probabilistic models or classifiers, which can be brittle under partial observability and limited interaction. Large language models (LLMs) offer a flexible alternative: by mapping short behavioral traces into high-level hypotheses, they can serve as world models over teammate behavior. We introduce \Collab, a language-based framework that classifies partner types using a behavior rubric derived from trajectory features, and extend it to \ReCollab, which incorporates retrieval-augmented generation (RAG) to stabilize inference with exemplar trajectories. In the cooperative Overcooked environment, \Collab effectively distinguishes teammate types, while \ReCollab consistently improves adaptation across layouts, achieving Pareto-optimal trade-offs between classification accuracy and episodic return. These findings demonstrate the potential of LLMs as behavioral world models for AHT and highlight the importance of retrieval grounding in challenging coordination settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。