arXiv:2507.20068cs.LGstat.ML2025-07被引 4

提出两种新方法,用辅助数据做强化学习评估时还能给出可信的置信区间。

PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data

  • 基于马尔可夫决策过程设计新保形预测方法,适用于连续状态空间。
  • 在库存、机器人、医疗等模拟器和MIMIC-IV真实数据上,置信区间覆盖率达标。
  • 适合高风险领域如医疗,需可靠评估不确定性的RL部署场景。

离策略评估(OPE)方法可在强化学习策略部署前估算其价值。近期研究发现,利用生成模型合成的辅助数据可提升OPE精度,但这类数据可能引入偏差,现有方法缺乏对数据增强下结果的严谨不确定性量化。在医疗等高风险领域,可靠的不确定性估计对安全部署至关重要。本文提出两种构建有效置信区间的OPE方法:第一种针对特定初始状态s下的策略值V^π(s),提出适用于连续状态空间马尔可夫决策过程的新保形预测方法;第二种面向所有初始状态的平均策略性能V^π,融合双重稳健估计与预测驱动推断思想。在涵盖库存管理、机器人、医疗及真实MIMIC-IV数据集的多个模拟器中,所提方法能有效利用辅助数据,并始终生成覆盖真实策略值的置信区间,优于以往方法。本工作为高风险领域中提供严格不确定性评估的OPE奠定了基础。

原文摘要 · Abstract (English)

Off-policy evaluation (OPE) methods estimate the value of a new reinforcement learning (RL) policy prior to deployment. Recent advances have shown that leveraging auxiliary datasets, such as those synthesized by generative models, can improve the accuracy of OPE methods. Unfortunately, such auxiliary datasets may also be biased, and existing methods for using data augmentation within OPE lack principled uncertainty quantification. In high stakes domains like healthcare, reliable uncertainty estimates are important for ensuring safe and informed deployment of RL policies. In this work, we propose two methods to construct valid confidence intervals for OPE with data augmentation. The first provides a confidence interval over $V^π(s)$, the policy value conditioned on an initial state $s$. To do so we introduce a new conformal prediction method suitable for Markov Decision Processes (MDPs) with continuous state spaces, extending prior work to higher-dimensional settings. Second, we consider the more common task of estimating the average policy performance over all initial states, $V^π$; we introduce a method that draws on ideas from doubly robust estimation and prediction powered inference. Across simulators spanning inventory management, robotics, healthcare, and a real healthcare dataset from MIMIC-IV, we find that our methods can effectively leverage auxiliary data and consistently produce confidence intervals that cover the ground truth policy values, unlike previously proposed methods. Our work enables a future in which OPE can provide rigorous uncertainty estimates for high-stakes domains.

强化学习置信区间医疗AIOPE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。