arXiv:2602.14914cs.LGcs.IR2026-02被引 2

最优加性基线比自归一化更优,显著降低离策略评估误差

Additive Control Variates Dominate Self-Normalisation in Off-Policy Evaluation

  • 用最优加性基线替代传统乘法控制变量
  • 理论证明新方法在均方误差上严格优于SNIPS
  • 适合研究推荐系统评估与离策略学习的学者

离策略评估(OPE)对无需高昂在线干预即可评估排序与推荐系统至关重要。自归一化逆倾向得分(SNIPS)是标准的方差缩减工具,采用乘法控制变量。近期离策略学习进展表明,加性控制变量(基线修正)可能表现更优,但缺乏理论保证。本文给出明确结论:我们证明,具有最优加性基线的β⋆-IPS估计器在均方误差上渐近主导SNIPS。通过解析分解方差差距,我们发现SNIPS渐近等价于使用一个特定但通常次优的加性基线。结果从理论上支持在排序与推荐系统中由自归一化转向最优基线修正。

原文摘要 · Abstract (English)

Off-policy evaluation (OPE) is essential for assessing ranking and recommendation systems without costly online interventions. Self-Normalised Inverse Propensity Scoring (SNIPS) is a standard tool for variance reduction in OPE, leveraging a multiplicative control variate. Recent advances in off-policy learning suggest that additive control variates (baseline corrections) may offer superior performance, yet theoretical guarantees for evaluation are lacking. This paper provides a definitive answer: we prove that $β^\star$-IPS, an estimator with an optimal additive baseline, asymptotically dominates SNIPS in Mean Squared Error. By analytically decomposing the variance gap, we show that SNIPS is asymptotically equivalent to using a specific -- but generally sub-optimal -- additive baseline. Our results theoretically justify shifting from self-normalisation to optimal baseline corrections for both ranking and recommendation.

离策略评估基线修正推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。