arXiv:2603.27952cs.IR2026-03

提出一种无需训练的推荐系统精度上限评估方法,可准确衡量推荐难度并指导数据选择。

On the Accuracy Limits of Sequential Recommender Systems: An Entropy-Based Approach

  • 基于熵理论构建无训练估算框架,不受候选集大小影响。
  • 在真实数据上与最优模型表现相关性高达0.914(Spearman rho)。
  • 可识别用户新颖性偏好差异,支持数据筛选与分群诊断。

序列推荐系统在离线精度上持续提升,但其性能距离数据固有的精度上限仍有多少差距尚不明确。可靠的、模型无关的精度天花板估计,有助于在模型开发前进行难度评估和潜力预判。现有预测性分析多结合熵估计与Fano不等式反演,但在推荐场景中易受候选集设定敏感性及低可预测性区域的缩放失真影响。本文提出一种基于熵的训练自由方法,用于量化序列推荐中的精度极限,实现候选集大小无关的估计。在受控合成生成器和多样真实基准上的实验表明,该估计器比基线更忠实反映理想难度,对候选集大小不敏感,并与当前顶尖序列推荐模型的离线最佳表现保持高度秩相关性(斯皮尔曼等级相关系数最高达0.914)。此外,该方法可通过用户分组(新颖性偏好、长尾暴露度、活跃度)揭示系统性可预测性差异,且高可预测性用户构成的训练集在数据受限时仍能获得优异下游性能。总体而言,该估计算法为评估可达成精度上限、支持用户群体诊断及数据驱动决策提供了实用工具。

原文摘要 · Abstract (English)

Sequential recommender systems have achieved steady gains in offline accuracy, yet it remains unclear how close current models are to the intrinsic accuracy limit imposed by the data. A reliable, model-agnostic estimate of this ceiling would enable principled difficulty assessment and headroom estimation before costly model development. Existing predictability analyses typically combine entropy estimation with Fano's inequality inversion; however, in recommendation they are hindered by sensitivity to candidate-space specification and distortion from Fano-based scaling in low-predictability regimes. We develop an entropy-induced, training-free approach for quantifying accuracy limits in sequential recommendation, yielding a candidate-size-agnostic estimate. Experiments on controlled synthetic generators and diverse real-world benchmarks show that the estimator tracks oracle-controlled difficulty more faithfully than baselines, remains insensitive to candidate-set size, and achieves high rank consistency with best-achieved offline accuracy across state-of-the-art sequential recommenders (Spearman rho up to 0.914). It also supports user-group diagnostics by stratifying users by novelty preference, long-tail exposure, and activity, revealing systematic predictability differences. Furthermore, predictability can guide training data selection: training sets constructed from high-predictability users yield strong downstream performance under reduced data budgets. Overall, the proposed estimator provides a practical reference for assessing attainable accuracy limits, supporting user-group diagnostics, and informing data-centric decisions in sequential recommendation.

推荐系统精度上限熵估计数据选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。