arXiv:2510.26026stat.MLcs.LG2025-10NeurIPS被引 2

为强化学习策略评估提供无需分布假设的不确定性量化方法

Conformal Prediction Beyond the Horizon: Distribution-Free Inference for Policy Evaluation

  • 结合分布强化学习与校准,用截断轨迹构造伪回报
  • 在山地车等环境上覆盖率达90%以上,优于传统方法
  • 适合高风险场景下需要可靠置信区间的策略评估

在高风险强化学习场景中,可靠的不确定性量化至关重要。本文提出一种统一的保形预测框架,用于无限时域策略评估,在在线和离线设置下均可构建无需分布假设的回报预测区间。方法融合分布强化学习与保形校准,解决未观测回报、时间依赖性和分布偏移等问题。提出基于截断轨迹的模块化伪回报构造方法,并采用基于经验回放和加权子采样的时间感知校准策略,缓解模型偏差并恢复近似可交换性,即使在策略转移下仍能实现不确定性量化。理论分析提供了考虑模型误设和重要性权重估计的覆盖率保证。实验结果表明,无论在合成环境还是山地车等基准环境中,该方法在覆盖率和可靠性上均显著优于标准分布强化学习基线。

原文摘要 · Abstract (English)

Reliable uncertainty quantification is crucial for reinforcement learning (RL) in high-stakes settings. We propose a unified conformal prediction framework for infinite-horizon policy evaluation that constructs distribution-free prediction intervals {for returns} in both on-policy and off-policy settings. Our method integrates distributional RL with conformal calibration, addressing challenges such as unobserved returns, temporal dependencies, and distributional shifts. We propose a modular pseudo-return construction based on truncated rollouts and a time-aware calibration strategy using experience replay and weighted subsampling. These innovations mitigate model bias and restore approximate exchangeability, enabling uncertainty quantification even under policy shifts. Our theoretical analysis provides coverage guarantees that account for model misspecification and importance weight estimation. Empirical results, including experiments in synthetic and benchmark environments like Mountain Car, show that our method significantly improves coverage and reliability over standard distributional RL baselines.

强化学习不确定性保形预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。