让大模型从被动回复转向主动协作,提升对话效率与用户满意度。
CollabLLM: From Passive Responders to Active Collaborators

- 通过多轮感知奖励模拟,评估回复的长期贡献,驱动模型主动挖掘用户意图。
- 在文档创作等任务上,任务完成率提升18.5%,交互性提高46.3%。
- 用户实测显示满意度升17.6%,耗时减少10.4%,适合复杂协作场景。
大型语言模型通常采用下一回合奖励进行训练,限制了其对长期交互的优化能力。因此,它们常对模糊或开放式的用户请求被动回应,无法帮助用户达成最终目标,导致对话效率低下。为解决这一问题,我们提出CollabLLM,一种新型通用训练框架,旨在增强多轮人机协作。其核心创新在于协作模拟机制,利用多轮感知奖励(Multiturn-aware Rewards)估算回复的长期贡献。通过强化学习微调这些奖励,CollabLLM超越了简单响应用户请求的层面,能够主动探索用户意图并提供有见地的建议,是迈向更以人为中心的AI的关键一步。我们还设计了一个包含三项挑战性任务的多轮交互基准,如文档创建。实验表明,CollabLLM在各项指标上显著优于基线模型:任务性能平均提升18.5%,由LLM评委评估的交互性提升46.3%。最后,我们进行了涵盖201名评委的大规模用户研究,结果显示,用户满意度提升17.6%,平均花费时间减少10.4%。
原文摘要 · Abstract (English)
Large Language Models are typically trained with next-turn rewards, limiting their ability to optimize for long-term interaction. As a result, they often respond passively to ambiguous or open-ended user requests, failing to help users reach their ultimate intents and leading to inefficient conversations. To address these limitations, we introduce CollabLLM, a novel and general training framework that enhances multiturn human-LLM collaboration. Its key innovation is a collaborative simulation that estimates the long-term contribution of responses using Multiturn-aware Rewards. By reinforcement fine-tuning these rewards, CollabLLM goes beyond responding to user requests, and actively uncovers user intent and offers insightful suggestions-a key step towards more human-centered AI. We also devise a multiturn interaction benchmark with three challenging tasks such as document creation. CollabLLM significantly outperforms our baselines with averages of 18.5% higher task performance and 46.3% improved interactivity by LLM judges. Finally, we conduct a large user study with 201 judges, where CollabLLM increases user satisfaction by 17.6% and reduces user spent time by 10.4%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。