评测大模型与用户互动中理解隐含需求的能力。
UserBench: An Interactive Gym Environment for User-Centric Agents
- 设计模拟用户逐步透露偏好的多轮交互环境
- 顶尖模型仅20%时间完全匹配用户意图,30%偏好未被发现
- 适合研究人机协作、具身智能与交互式推理的学者
基于大语言模型的智能体在推理和工具使用方面取得了显著进展,但在目标模糊、动态变化或间接表达时主动与用户协作的能力仍不充分。为弥补这一空白,我们提出UserBench——一个以用户为中心的基准测试,用于评估智能体在多轮、偏好驱动交互中的表现。该环境包含模拟用户,初始目标不明确,并逐步揭示偏好,要求智能体主动澄清意图并使用工具做出合理决策。对主流开源与闭源大模型的评估显示,任务完成率与用户一致性之间存在显著脱节:模型在平均情况下仅20%的时间提供与所有用户意图完全一致的答案,即使最先进的模型通过主动交互也未能发现超过30%的用户偏好。这些结果凸显了构建真正协同伙伴型智能体的挑战。UserBench提供了一个可交互的评估环境,用于测量并推动该关键能力的发展。
原文摘要 · Abstract (English)
Large Language Models (LLMs)-based agents have made impressive progress in reasoning and tool use, enabling them to solve complex tasks. However, their ability to proactively collaborate with users, especially when goals are vague, evolving, or indirectly expressed, remains underexplored. To address this gap, we introduce UserBench, a user-centric benchmark designed to evaluate agents in multi-turn, preference-driven interactions. UserBench features simulated users who start with underspecified goals and reveal preferences incrementally, requiring agents to proactively clarify intent and make grounded decisions with tools. Our evaluation of leading open- and closed-source LLMs reveals a significant disconnect between task completion and user alignment. For instance, models provide answers that fully align with all user intents only 20% of the time on average, and even the most advanced models uncover fewer than 30% of all user preferences through active interaction. These results highlight the challenges of building agents that are not just capable task executors, but true collaborative partners. UserBench offers an interactive environment to measure and advance this critical capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。