arXiv:2409.04810cs.IR2024-09中稿 · WWW'2025被引 1

现有推荐模型评估方法在随机暴露数据上不可靠,新方案可更准确衡量去偏效果。

Debias Can be Unreliable: Mitigating Bias Issue in Evaluating Debiasing Recommendation

  • 提出无偏召回评估(URE)方法,修正随机暴露数据的偏差问题。
  • 实验证明传统方法在随机暴露数据上得出的召回率与真实值不一致。
  • 适用于评估去偏推荐系统,尤其对真实场景数据研究者有参考价值。

近期研究通过引入去偏方法显著提升了推荐模型性能。由于全曝光数据难以获取,现有方法普遍采用随机暴露数据作为代理,使用传统评估方案衡量推荐表现。然而本文揭示,传统评估方案在随机暴露数据上并不适用,导致基于随机暴露数据获得的召回率与全曝光数据下的真实召回率存在不一致。这种不一致性表明此前去偏技术实验结论可能存在可靠性问题,亟需在随机暴露数据上实现无偏召回评估。为此,我们提出无偏召回评估(URE)方案,通过调整随机暴露数据的使用方式,无偏估计全曝光数据上的真实召回性能。我们提供了理论依据证明URE的合理性,并在多个真实世界数据集上进行了广泛实验,验证其有效性。

原文摘要 · Abstract (English)

Recent work has improved recommendation models remarkably by equipping them with debiasing methods. Due to the unavailability of fully-exposed datasets, most existing approaches resort to randomly-exposed datasets as a proxy for evaluating debiased models, employing traditional evaluation scheme to represent the recommendation performance. However, in this study, we reveal that traditional evaluation scheme is not suitable for randomly-exposed datasets, leading to inconsistency between the Recall performance obtained using randomly-exposed datasets and that obtained using fully-exposed datasets. Such inconsistency indicates the potential unreliability of experiment conclusions on previous debiasing techniques and calls for unbiased Recall evaluation using randomly-exposed datasets. To bridge the gap, we propose the Unbiased Recall Evaluation (URE) scheme, which adjusts the utilization of randomly-exposed datasets to unbiasedly estimate the true Recall performance on fully-exposed datasets. We provide theoretical evidence to demonstrate the rationality of URE and perform extensive experiments on real-world datasets to validate its soundness.

推荐系统去偏评估召回率数据偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。