arXiv:2603.21485cs.LG2026-03被引 1

提出新方法在确定性日志下实现低偏差排序策略评估

Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies

  • 用用户点击概率替代日志策略随机性做重要性加权
  • 在完全确定的日志策略下,偏差显著低于现有方法
  • 适合需要真实数据评估排序算法的工业场景

离线策略评估(OPE)是算法排序系统中的关键问题,目标是仅通过不同日志策略收集的离线数据,估计新排序策略的期望性能。现有估计器如排名级和位置级逆倾向得分(IPS)需依赖日志策略的充分随机性,当日志策略完全确定时会引入严重偏差。本文提出新型估计器——基于点击的逆倾向得分(CIPS),利用用户点击行为的内在随机性,将点击概率作为新的重要性权重,即使在完全确定的日志策略下也能实现低偏差评估。我们提供了所提估计器的偏差与方差理论分析,并通过合成数据与真实世界实验验证,在多种完全确定日志设置下,其偏差显著优于强基线方法。

原文摘要 · Abstract (English)

Off-Policy Evaluation (OPE) is an important practical problem in algorithmic ranking systems, where the goal is to estimate the expected performance of a new ranking policy using only offline logged data collected under a different, logging policy. Existing estimators, such as the ranking-wise and position-wise inverse propensity score (IPS) estimators, require the data collection policy to be sufficiently stochastic and suffer from severe bias when the logging policy is fully deterministic. In this paper, we propose novel estimators, Click-based Inverse Propensity Score (CIPS), exploiting the intrinsic stochasticity of user click behavior to address this challenge. Unlike existing methods that rely on the stochasticity of the logging policy, our approach uses click probability as a new form of importance weighting, enabling low-bias OPE even under deterministic logging policies where existing methods incur substantial bias. We provide theoretical analyses of the bias and variance properties of the proposed estimators and show, through synthetic and real-world experiments, that our estimators achieve significantly lower bias compared to strong baselines, for a range of experimental settings with completely deterministic logging policies.

离线评估排序算法逆倾向得分点击建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。