首个基于真实对话的LLM个性化评测基准,助力AI更懂用户。
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
- 从真实对话中构建2500组长期交互数据,含人工验证的记忆结构。
- 发现主流模型在提取隐性用户特征、更新记忆等方面存在明显短板。
- 适合研究个性化、记忆机制或人机交互的学者与开发者使用。
随着大语言模型演变为终身陪伴式AI助理,个性化成为关键挑战。当前进展受限于缺乏金标准评测基准:现有测评要么忽略个性化信息管理,要么依赖合成对话,与真实对话存在分布差异。为此,我们提出AlpsBench,一个基于真实人类-大模型对话(WildChat)构建的个性化评测基准。该基准包含2,500条长期交互序列,配有经人工验证的结构化记忆,涵盖显性和隐性个性化信号。我们定义了四项核心任务——个性化信息提取、更新、检索与利用,并建立完整的记忆管理生命周期评估协议。对前沿大模型与记忆导向系统进行评测发现:(i) 模型难以可靠提取用户潜在特质;(ii) 即使最强模型在记忆更新上也面临性能瓶颈;(iii) 面对大量干扰项时,检索准确率急剧下降;(iv) 显式记忆机制虽提升召回率,但并不天然带来更符合偏好或情感共鸣的回应。AlpsBench旨在提供一个全面的评估框架。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) evolve into lifelong AI assistants, LLM personalization has become a critical frontier. However, progress is currently bottlenecked by the absence of a gold-standard evaluation benchmark. Existing benchmarks either overlook personalized information management that is critical for personalization or rely heavily on synthetic dialogues, which exhibit an inherent distribution gap from real-world dialogue. To bridge this gap, we introduce AlpsBench, An LLM PerSonalization benchmark derived from real-world human-LLM dialogues. AlpsBench comprises 2,500 long-term interaction sequences curated from WildChat, paired with human-verified structured memories that encapsulate both explicit and implicit personalization signals. We define four pivotal tasks - personalized information extraction, updating, retrieval, and utilization - and establish protocols to evaluate the entire lifecycle of memory management. Our benchmarking of frontier LLMs and memory-centric systems reveals that: (i) models struggle to reliably extract latent user traits; (ii) memory updating faces a performance ceiling even in the strongest models; (iii) retrieval accuracy declines sharply in the presence of large distractor pools; and (iv) while explicit memory mechanisms improve recall, they do not inherently guarantee more preference-aligned or emotionally resonant responses. AlpsBench aims to provide a comprehensive framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。