解决推荐系统隐式反馈中的曝光偏差问题,提升模型鲁棒性与评估稳定性。
Counterfactual Risk Minimization with IPS-Weighted BPR and Self-Normalized Evaluation in Recommender Systems
- 用加权BPR+倾向性正则化,降低极端权重带来的方差放大。
- 在合成数据和MovieLens 100K上验证,泛化能力更强且评估方差更低。
- 适合做真实场景下基于日志数据的推荐系统训练与评估。
从记录的隐式反馈中学习和评估推荐系统面临曝光偏差挑战。虽然逆倾向评分(IPS)可纠正此偏差,但常导致高方差和不稳定性。本文提出一个简单有效的流程:将IPS加权训练与增强倾向性正则化的IPS加权贝叶斯个性化排序(BPR)目标结合。比较了直接法(DM)、IPS与自归一化IPS(SNIPS)在离线策略评估中的表现,证明IPS加权训练能提升模型在有偏曝光下的鲁棒性。所提倾向性正则化进一步缓解极端倾向权重引发的方差放大,实现更稳定的估计。在合成数据和MovieLens 100K上的实验表明,该方法在无偏曝光下泛化性能更好,同时相比朴素法与标准IPS方法显著降低评估方差,为真实推荐场景中的反事实学习与评估提供实用指导。
原文摘要 · Abstract (English)
Learning and evaluating recommender systems from logged implicit feedback is challenging due to exposure bias. While inverse propensity scoring (IPS) corrects this bias, it often suffers from high variance and instability. In this paper, we present a simple and effective pipeline that integrates IPS-weighted training with an IPS-weighted Bayesian Personalized Ranking (BPR) objective augmented by a Propensity Regularizer (PR). We compare Direct Method (DM), IPS, and Self-Normalized IPS (SNIPS) for offline policy evaluation, and demonstrate how IPS-weighted training improves model robustness under biased exposure. The proposed PR further mitigates variance amplification from extreme propensity weights, leading to more stable estimates. Experiments on synthetic and MovieLens 100K data show that our approach generalizes better under unbiased exposure while reducing evaluation variance compared to naive and standard IPS methods, offering practical guidance for counterfactual learning and evaluation in real-world recommendation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。