arXiv:2507.06961stat.MLcs.LG2025-07ICML被引 2

解决离线策略评估中缺失数据导致的偏差问题

Off-Policy Evaluation Under Nonignorable Missing Data

  • 提出加权估计算法应对非忽略式缺失数据
  • 证明忽略缺失数据时估计无偏,非忽略则有偏
  • 适合实际应用中存在数据缺失的强化学习研究者

离线策略评估(OPE)旨在利用从不同策略收集的离线数据估计目标策略的价值。然而,在真实场景中,记录数据常存在缺失。尽管OPE已广泛研究,但缺失数据对评估结果的影响尚缺乏理论认知。本文研究单调缺失情况下的OPE,理论上证明:在可忽略缺失下价值估计无偏,而在不可忽略(信息性)缺失下会产生偏差。为保持估计一致性,我们提出逆概率加权价值估计器,并进行统计推断以量化不确定性。一系列数值实验表明,所提方法在缺失数据下能提供更可靠的值推理。

原文摘要 · Abstract (English)

Off-Policy Evaluation (OPE) aims to estimate the value of a target policy using offline data collected from potentially different policies. In real-world applications, however, logged data often suffers from missingness. While OPE has been extensively studied in the literature, a theoretical understanding of how missing data affects OPE results remains unclear. In this paper, we investigate OPE in the presence of monotone missingness and theoretically demonstrate that the value estimates remain unbiased under ignorable missingness but can be biased under nonignorable (informative) missingness. To retain the consistency of value estimation, we propose an inverse probability weighted value estimator and conduct statistical inference to quantify the uncertainty of the estimates. Through a series of numerical experiments, we empirically demonstrate that our proposed estimator yields a more reliable value inference under missing data.

强化学习离线评估缺失数据统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。