arXiv:2604.26996cs.IR2026-04被引 1

提出新指标与测试集,发现现有系统忽视情境相关性。

LUCid: Redefining Relevance For Lifelong Personalization

  • 引入PA分数量化语义相近性偏差,区分情境与语义相关性。
  • 在最难任务上检索召回率趋近于零,生成结果对齐率仅约50%。
  • 适合关注长期个性化与真实场景适应性的研究者参考。

现有工作主要依赖语义相近性来识别长期个性化中的相关内容,但情境相关性对用户实际任务和上下文更为关键。本文提出邻近优势(PA)评分,用于量化语义相近性偏差,并指出当前个性化基准大多混淆了语义与情境相近性,无法判断系统是否真正捕捉情境相关性。为此,我们构建LUCid诊断基准,包含1,936个用户查询及其长期交互历史,旨在将情境相关性从语义相近性中分离。在现代个性化流程的各个阶段(检索、重排序、生成)进行实验,结果显示显著性能下降:最困难实例的检索召回率趋近于零,即使最先进的模型如Gemini-3-Flash、GPT-5.4和Claude Haiku,响应对齐率也仅维持在约50%,暴露出当前系统所编码的相关性与长期个性化需求之间存在根本性错配。

原文摘要 · Abstract (English)

Work to date has mainly relied on semantic proximity to identify relevant content for lifelong personalization. However, situational relevance is often more important for determining which information is useful for a user's actual task and context. In this paper, we introduce the Proximity Advantage (PA) score, a metric for quantifying semantic proximity bias, and show that existing personalization benchmarks largely conflate semantic and situational proximity, leaving it unclear whether current systems truly capture situational relevance. To support this metric, we introduce LUCid, a diagnostic benchmark of 1,936 user queries paired with long interaction histories, designed to isolate situational relevance from semantic proximity. Our experiments across different stages of the modern personalization pipeline (retrieval, reranking, and generation) reveal significant performance collapse: retrieval recall drops to near zero on the hardest instances, and response alignment remains near 50\% even for state-of-the-art models such as Gemini-3-Flash, GPT-5.4, and Claude Haiku, highlighting a fundamental mismatch between the relevance encoded by current systems and what lifelong personalization demands.

个性化情境感知评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。