arXiv:2608.29600cs.LGcs.IR2026-09

提出自适应双重稳健方法,提升排序策略离线评估精度。

Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior

论文配图:Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior
图 1 · 摘自论文原文
  • 结合自适应重要性加权与奖励回归进行控制变量修正
  • 在真实用户行为模型已知时保证无偏,且方差低于AIPS
  • 适用于长排序和小样本场景,适合推荐系统评估研究者

排序策略的离线评估(OPE)面临挑战:从候选集选择并排序多个项目,导致可能的排序组合数随候选数量和排序长度呈组合爆炸。传统逆倾向得分(IPS)因需计算完整排序概率比,方差过大。独立IPS(IIPS)与奖励交互IPS(RIPS)通过固定用户浏览假设降低方差,但当假设不匹配实际行为时会引入偏差。自适应逆倾向得分(AIPS)通过自适应边缘化影响各位置奖励的动作重要性权重,在已知真实用户行为模型时,可达到该类无偏IPS估计器的最小方差。然而,其在长排序下估计精度仍可能下降,且未使用奖励模型进行残差修正。本文提出自适应双重稳健(ADR),将自适应重要性加权与奖励回归结合,通过控制变量修正实现残差校正。我们证明了当真实用户行为模型已知时,ADR具有无偏性,并给出了其相对于AIPS降低方差的充分条件。在每种条件下10,000次模拟实验中,ADR在不同日志数据规模和排序长度下均优于AIPS及传统排序OPE估计器的均方误差。

原文摘要 · Abstract (English)

Off-policy evaluation (OPE) of ranking policies is challenging be- cause selecting and ordering multiple items from a candidate set makes the number of possible rankings grow combinatorially with the number of candidates and the ranking length. Consequently, Inverse Propensity Scoring (IPS), whose importance weight is the full-ranking probability ratio under the evaluation and logging policies, can have excessive variance. Independent IPS (IIPS) and Reward Interaction IPS (RIPS) reduce variance by imposing fixed assumptions on how users browse rankings, but may introduce bias when those assumptions mismatch actual behavior. Adaptive Inverse Propensity Scoring (AIPS) addresses this trade-off by adap- tively marginalizing importance weights over the actions that affect each position-wise reward. It attains minimum variance within a class of unbiased IPS-based estimators when the true user be- havior model is observed. However, its estimation accuracy may still degrade for longer rankings, and AIPS does not use a reward model for residual correction. We propose Adaptive Doubly Robust (ADR), which combines adaptive importance weighting with re- ward regression through a control-variate correction. We establish its unbiasedness when the true user behavior model is observed and characterize a sufficient condition under which it reduces vari- ance relative to AIPS. Across synthetic experiments with 10,000 simulations per condition, ADR improves mean squared error over AIPS and conventional ranking OPE estimators across a range of logged-data sizes and ranking lengths.

离线评估排序优化双稳健估计推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。