构建长时个性化偏好评估基准,揭示大模型适应用户长期习惯的瓶颈
Towards Natural Personalization: Evaluating Long-Horizon Preference Following in Personalized User-LLM Interactions
- 设计包含100个虚拟用户、1300条偏好的真实交互评估框架
- 发现上下文越长、表达越隐晦,模型表现越差,泛化能力仍不足
- 适合研究个性化助手、长程对话理解与用户建模的学者参考
大语言模型正越来越多地作为个人助理,用户在长期互动中可能持续传递个性化偏好。然而,评估模型在自然、长期场景下对这些偏好的遵循能力仍不充分。本文提出RealPref,一个用于评估个性化用户-大模型交互中自然偏好跟随的基准。该基准包含100个合成用户档案、1300条个性化偏好,涵盖4类偏好表达方式(从明确到隐含),以及长时序交互历史。测试任务包括多项选择、是非判断和开放问答三类,并提供细粒度评分标准供模型作为评判者使用。结果表明,随着上下文长度增加和偏好表达趋于隐含,模型性能显著下降,且在未见场景中泛化理解用户偏好仍存在挑战。RealPref及实验发现为未来开发更懂用户的智能助手提供了基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly serving as personal assistants, where users may share individual preferences over extended interactions. However, assessing how well LLMs can follow these preferences in natural, long-term situations remains underexplored. This work proposes RealPref, a benchmark for evaluating natural preference-following in personalized user-LLM interactions. RealPref features 100 synthetic user profiles, 1300 personalized preferences, 4 types of preference expression (from explicit to implicit), and long-horizon interaction histories. It explored three types of test tasks (multiple-choice, true-or-false, and open-ended), with granular rubrics for LLM-as-a-judge evaluation. Results indicate that LLM performance drops significantly as context length grows and preference expression becomes more implicit, and that generalizing user preference understanding to unseen scenarios poses further challenges. RealPref and these findings provide a foundation for future research to develop user-aware LLM assistants that better adapt to individual needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。