arXiv:2604.17283cs.CLcs.AI2026-04被引 10

构建6个月用户偏好演变数据集,揭示长时个性化核心瓶颈

HorizonBench: Long-Horizon Personalization with Evolving Preferences

论文配图:HorizonBench: Long-Horizon Personalization with Evolving Preferences
图 1 · 摘自论文原文
  • 基于心理状态图生成模拟对话,实现偏好变化的精准溯源
  • 25个前沿模型平均表现仅达20%随机基线,最优仅52.8%
  • 多数模型无法更新用户状态,30%以上仍沿用过时偏好

用户偏好随数月交互持续演变,追踪需识别由生活事件引发的偏好变更。本文定义该问题为长时个性化,并指出当前进展受限于数据稀缺与测量困难——缺乏自然交互数据与真实偏好转折的标注。为此,我们设计数据生成器,基于结构化心理状态图生成对话,为6个月时间跨度内的每项偏好变化提供真实来源。由此构建HorizonBench基准,包含360名模拟用户的4,245条数据,平均对话轮次约4,300,总词元数约163K。该基准可用于长上下文建模、记忆增强架构、心智理论推理与用户建模评估。在25个前沿模型测试中,最佳模型准确率为52.8%,多数模型表现不高于20%随机基线。当模型在演化偏好上出错时,超过三分之一情况下仍选择用户初始声明值,未追踪其最新状态。该信念更新失败现象在不同上下文长度与表达清晰度下均存在,表明状态追踪能力是长时个性化的核心瓶颈。

原文摘要 · Abstract (English)

User preferences evolve across months of interaction, and tracking them requires inferring when a stated preference has been changed by a subsequent life event. We define this problem as long-horizon personalization and observe that progress on it is limited by data availability and measurement, with no existing resource providing both naturalistic long-horizon interactions and the ground-truth provenance needed to diagnose why models fail. We introduce a data generator that produces conversations from a structured mental state graph, yielding ground-truth provenance for every preference change across 6-month timelines, and from it construct HorizonBench, a benchmark of 4,245 items from 360 simulated users with 6-month conversation histories averaging ~4,300 turns and ~163K tokens. HorizonBench provides a testbed for long-context modeling, memory-augmented architectures, theory-of-mind reasoning, and user modeling. Across 25 frontier models, the best model reaches 52.8% and most score at or below the 20% chance baseline. When these models err on evolved preferences, over a third of the time they select the user's originally stated value without tracking the updated user state. This belief-update failure persists across context lengths and expression explicitness levels, identifying state-tracking capability as the primary bottleneck for long-horizon personalization.

长时个性化偏好演化基准测试状态追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。