评测大模型在长期交互中理解并主动适应用户偏好的能力
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

- 构建时序化用户任务序列,模拟真实碎片化交互
- 顶尖模型仍难以有效捕捉和更新用户偏好,表现差距明显
- 适合研究个性化智能体与人机长期协作的学者
大型语言模型已发展为可与用户协作完成现实任务的智能体。有效协作依赖于理解用户未明说的需求,而用户意图常隐含于零散的日常交互中,需兼具个性化建模与主动交互能力。现有评测主要关注推理与工具使用,忽视了在真实场景中推断和利用用户偏好的挑战。为此,我们提出VitaBench 2.0,一个评估个性化与主动型智能体在长期用户交互中表现的基准。任务以时序序列组织,用户偏好嵌入于碎片化、异构的交互中,成功完成任务需智能体持续从交互中提取、利用并更新偏好。通过要求智能体识别缺失信息并主动向用户或环境获取,进一步评估其主动性。我们提供可扩展的记忆接口,支持不同记忆架构的可控对比。对一系列前沿专有及开源大模型进行评测,结果表明即使最先进的模型在真实个性化任务上仍面临巨大挑战,揭示了当前能力与实际需求间的显著差距。深入分析还揭示了现有智能体在真实个性化决策中的失败模式与能力瓶颈,为未来模型改进提供洞见。
原文摘要 · Abstract (English)
Large language models (LLMs) have evolved into interactive agents that collaborate with users in real-world tasks. Effective collaboration in such settings increasingly depends on understanding the user beyond what is explicitly stated, as user intent is often reflected in fragmented daily interactions and requires both personalized modeling and proactive interaction. However, existing agent benchmarks primarily evaluate reasoning and tool use, largely overlooking the challenges of inferring and leveraging user preferences in realistic scenarios. To address this gap, we introduce VitaBench 2.0, a benchmark for evaluating personalized and proactive agent behavior in long-term user interactions. In VitaBench 2.0, tasks are organized as temporally ordered sequences for individual users, where preferences are embedded in fragmented and heterogeneous interactions. Successful completion of tasks requires the agent to continuously extract, utilize, and update user preferences from these interactions. We further evaluate proactiveness through tasks that require agents to recognize missing information and actively acquire it from users or environments before making decisions. To support systematic analysis, we provide an extensible memory interface that enables controlled comparison across different memory architectures. We benchmark a diverse set of frontier proprietary and open-source LLMs. Results show that real-world personalization remains highly challenging even for state-of-the-art models, revealing a substantial gap between current capabilities and practical requirements. Extensive analysis further reveals the failure modes and capability bottlenecks of current agents in real-world personalized decision-making, providing insights for future model improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。