arXiv:2606.20206stat.MLcs.LG2026-06中稿 · ICML

解决奖励缺失非随机下的离线评估难题,提升医疗等场景的策略评估准确性。

Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at Random

论文配图:Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at Random
图 1 · 摘自论文原文
  • 引入未来状态作为影子变量,识别完整数据下的条件奖励均值。
  • 设计桥接函数恢复条件奖励,避免显式建模缺失机制。
  • 适用于医疗等缺失机制复杂的场景,尤其适合关注历史缺失信息的策略评估。

在离线强化学习中,由于记录稀疏或奖励被截断,日志数据中的即时奖励常缺失。这一问题在医疗和营销等领域尤为突出。本文研究有限时域马尔可夫决策过程下奖励缺失非随机(MNAR)时的离线策略评估(OPE)。MNAR破坏了可忽略性,即使在给定状态与动作后仍存在选择偏差。为此,我们形式化了一个依赖奖励的倾向性模型,并利用未来状态作为影子变量来识别全数据条件均值奖励。进一步提出桥接函数,无需显式建模MNAR机制即可恢复条件均值奖励,并通过极小极大过程估计,避免双重采样。基于上述识别结果,我们提出一种类拟合Q-评估的估计器,可传播恢复后的奖励,且允许目标策略依赖于过去的缺失指示。最后,建立了该估计器的一致性及有限样本误差界,并在模拟数据和MIMIC-III脓毒症数据上验证了方法优于现有方法的性能。

原文摘要 · Abstract (English)

In offline Reinforcement Learning, immediate rewards in logged batch data are often unobserved due to sparse or irregular record-keeping, or censored beyond certain reward values. This issue arises in practical settings, including health care and marketing. We investigate off-policy evaluation (OPE) in finite-horizon Markov decision processes when rewards are missing not at random (MNAR), which breaks ignorability and induces selection bias even after conditioning on states and actions. To address this, we formalize a reward-dependent propensity model and use future states as shadow variables to identify the full-data conditional mean reward. We further introduce a bridge function that recovers the conditional mean reward without explicitly modeling the MNAR mechanism, and estimate it via a min-max procedure to avoid double sampling. Building upon these identification results, we propose an Fitted-Q-Evaluation-style estimator that propagates the recovered rewards while allowing target policies to depend on past missingness indicators. Finally, we establish consistency and finite-sample error bounds for our OPE estimator, and show through experiments the strong performance of our method compared to existing methods on simulated and MIMIC-III Sepsis data.

离线评估缺失数据强化学习医疗应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。