提出新方法评估推荐系统效果,显著降低误差。
Off-Policy Evaluation of Ranking Policies via Embedding-Space User Behavior Modeling
- 基于嵌入空间建模用户行为,改进评估偏差。
- 实验显示新方法均方误差最低,优于现有方法。
- 适合大规模推荐系统离线评估,尤其在复杂场景下。
在排名动作空间庞大的推荐场景中,仅用历史带宽数据评估新推荐策略的离线策略评估(OPE)至关重要。为解决现有估计器方差过高的问题,本文引入两个新假设:无直接排名影响,以及在排名嵌入空间上的用户行为模型。在此基础上,提出广义边际逆倾向得分(GMIPS)估计器,具有更优的统计性质。实验表明,GMIPS实现最低均方误差(MSE)。其中,边际奖励交互IPS(MRIPS)采用基于级联行为假设的双重边际化重要性权重,在排名空间扩大且假设不成立时仍能有效平衡偏差与方差,表现稳定。
原文摘要 · Abstract (English)
Off-policy evaluation (OPE) in ranking settings with large ranking action spaces, which stems from an increase in both the number of unique actions and length of the ranking, is essential for assessing new recommender policies using only logged bandit data from previous versions. To address the high variance issues associated with existing estimators, we introduce two new assumptions: no direct effect on rankings and user behavior model on ranking embedding spaces. We then propose the generalized marginalized inverse propensity score (GMIPS) estimator with statistically desirable properties compared to existing ones. Finally, we demonstrate that the GMIPS achieves the lowest MSE. Notably, among GMIPS variants, the marginalized reward interaction IPS (MRIPS) incorporates a doubly marginalized importance weight based on a cascade behavior assumption on ranking embeddings. MRIPS effectively balances the trade-off between bias and variance, even as the ranking action spaces increase and the above assumptions may not hold, as evidenced by our experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。