构建多人协作游戏评测框架,提升大模型与多样人类角色的协同能力。
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement

- 通过模拟不同性格玩家行为,构建真实协作环境。
- 训练模型效率提升19.5%,情感适配性提高24.4%。
- 适合研究人机协作、智能体交互与社会性建模的学者。
尽管基于大模型的智能体在单任务中表现优异,但与真实人类伙伴进行有效协作仍具挑战。现有对话级协作研究缺乏实际互动与行为执行,亟需能实现情境化、沉浸式协作的协作游戏环境。为此,本文提出 CollabBench,一个用于评估与训练协作智能体的基准平台。该平台包含多样玩家行为模拟流程,以及通过智能体回溯统一推理、沟通与行动的协作训练范式,并采用混合奖励机制平衡任务效率与情感适应性。我们进一步将经典环境扩展为 CWAH-MultiPlayer 与 Cook-MultiPlayer,以在多类型人格下系统评估。实验显示,训练模型在效率与情感指标上均优于基线模型,分别提升19.5%与24.4%。深入分析揭示了现有模型的协作瓶颈,为未来协作训练提供方向。
原文摘要 · Abstract (English)
While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conversation-level collaborative studies lack grounded interaction and behavioral execution, motivating the need for cooperative game environments that enable contextualized and immersive collaboration. To this end, this paper proposes CollabBench, a benchmark for evaluating and training collaborative agents in cooperative games. CollabBench features a Diverse Player Profile Simulation pipeline to model varied players behaviors, and a Collaborative Agentic Training paradigm that unifies reasoning, communication, and action via agentic rollouts, optimized with a hybrid reward balancing task efficiency and affective adaptation. We further extend classic environments to CWAH-MultiPlayer and Cook-MultiPlayer for systematic evaluation under diverse personalities. Experiments with efficiency and affective metrics show that our trained models outperform base models, achieving 19.5% higher efficiency and 24.4% improved affective performance. Further analysis reveals key collaborative limitations of existing models and offers insights for future collaborative training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。