解决推荐系统中缺失数据带来的偏差问题,提升评估准确性。
Off-Policy Evaluation for Recommendations with Missing-Not-At-Random Rewards
- 用日志策略和奖励观测概率作为倾向得分,联合建模双重偏差
- 在奖励缺失非随机时仍保持低偏差,性能优于现有方法
- 适合处理真实场景中复杂数据缺失的推荐评估任务
无偏推荐学习(URL)和离策略评估/学习(OPE/L)技术能有效缓解显示位置与记录策略带来的数据偏差,从而持续提升推荐性能。然而,当日志数据中同时存在这两种偏差时,现有估计器可能产生显著偏差。本文首次分析了奖励缺失非随机情况下OPE估计器的位置偏差。为缓解双重偏差,提出一种新估计器,利用两个日志策略与奖励观测的概率作为倾向得分。实验表明,该估计器在奖励观测偏差加剧时仍表现更优,显著优于其他对比方法。
原文摘要 · Abstract (English)
Unbiased recommender learning (URL) and off-policy evaluation/learning (OPE/L) techniques are effective in addressing the data bias caused by display position and logging policies, thereby consistently improving the performance of recommendations. However, when both bias exits in the logged data, these estimators may suffer from significant bias. In this study, we first analyze the position bias of the OPE estimator when rewards are missing not at random. To mitigate both biases, we propose a novel estimator that leverages two probabilities of logging policies and reward observations as propensity scores. Our experiments demonstrate that the proposed estimator achieves superior performance compared to other estimators, even as the levels of bias in reward observations increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。