用认知模型模拟用户长期行为,评估助理的个性化能力
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
- 基于BDI模型构建用户认知框架,生成连贯人生轨迹
- 覆盖8个生活领域1200个场景,支持多轮交互评估
- 揭示当前大模型在隐性意图和长期偏好建模上的不足
大语言模型的快速发展推动了通用人工智能助手的发展。然而,现有个性化助手评测基准与真实用户-助手交互存在偏差,难以捕捉外部环境复杂性和用户认知状态。为此,我们提出LifeSim,一种通过信念-欲望-意图(BDI)模型在物理环境中建模用户认知、生成连贯人生轨迹并模拟意图驱动交互行为的用户模拟器。基于LifeSim,我们构建LifeSim-Eval,一个涵盖8个生活领域、1200种多样化场景的综合性长周期个性化辅助评测基准。该基准采用多轮交互方式,评估模型完成显性与隐性意图、恢复用户画像及生成高质量响应的能力。在单场景与长周期设置下,实验表明当前大模型在处理隐性意图和长期用户偏好建模方面存在显著局限。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has accelerated progress toward universal AI assistants. However, existing benchmarks for personalized assistants remain misaligned with real-world user-assistant interactions, failing to capture the complexity of external contexts and users' cognitive states. To bridge this gap, we propose LifeSim, a user simulator that models user cognition through the Belief-Desire-Intention (BDI) model within physical environments for coherent life trajectories generation, and simulates intention-driven user interactive behaviors. Based on LifeSim, we introduce LifeSim-Eval, a comprehensive benchmark for multi-scenario, long-horizon personalized assistance. LifeSim-Eval covers 8 life domains and 1,200 diverse scenarios, and adopts a multi-turn interactive method to assess models' abilities to complete explicit and implicit intentions, recover user profiles, and produce high-quality responses. Under both single-scenario and long-horizon settings, our experiments reveal that current LLMs face significant limitations in handling implicit intention and long-term user preference modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。