arXiv:2512.23694stat.MLcs.LG2025-12被引 2

提出新方法提升离线强化学习中长期价值预测的可靠性。

Bellman Calibration for $V$-Learning in Offline Reinforcement Learning

  • 引入贝尔曼校准,通过比较预测值与目标值的一致性来诊断误差。
  • 在无贝尔曼完备性假设下,实现非参数率的误差控制。
  • 适用于任何已有价值函数模型,可后处理提升精度。

离线强化学习中长期价值预测难以可靠,因拟合值函数方法同时涉及自举、函数逼近和分布偏移,而标准保证通常依赖贝尔曼完备性或可实现性。本文提出贝尔曼校准——一种弱可靠性准则,要求相似预测值的状态其平均贝尔曼目标应与预测一致。该准则生成标量校准误差,可基于离策略数据通过双重稳健的贝尔曼目标估计获得。我们进一步提出迭代贝尔曼校准,一种模型无关的后处理方法,通过拟合原预测的一维映射(直方图与保序变体)进行校正。理论证明在有限样本下,校准误差以一维非参数速率受控,无需贝尔曼完备性或值函数可实现性。值误差界分离了统计估计、有限迭代与近似误差,明确校准何时提升预测性能,以及何时受限于原始预测器的信息或覆盖不足。

原文摘要 · Abstract (English)

Reliable long-horizon value prediction is difficult in offline reinforcement learning because fitted value methods combine bootstrapping, function approximation, and distribution shift, while standard guarantees often require Bellman completeness or realizability. We introduce Bellman calibration, a weak reliability criterion requiring that states assigned similar predicted values have average Bellman targets that agree with those predictions. This criterion yields a scalar calibration error for diagnosing systematic numerical miscalibration, which we estimate from off-policy data using doubly robust Bellman target estimates. We then propose Iterated Bellman Calibration, a model-agnostic post-hoc procedure that recalibrates any learned value predictor by fitting a one-dimensional map of its original prediction, with histogram and isotonic variants. We prove finite-sample guarantees showing that Bellman calibration error is controlled at one-dimensional nonparametric rates without Bellman completeness or value-function realizability. Our value-error bounds separate statistical estimation, finite-iteration, and approximation errors, clarifying when calibration improves value prediction and when its gains are limited by the information in the original predictor or insufficient coverage.

强化学习离线学习价值函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。