提出一种后处理方法,让离线强化学习的策略评估更准确。
Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning
- 通过单调修正初始比率估计,修复状态-动作分布偏差
- 保证校准后估计与真实目标分布的误差在统计可接受范围内
- 适合需要高精度策略评估的研究者,尤其关注离线学习
边际重要性加权通过折扣占用率重加权离线状态-动作样本来评估目标策略,其特性由伴随贝尔曼方程刻画。现有极小极大、原对偶及拟合固定点估计器因函数类近似、正则化或优化不完全,可能遗留占用平衡残差,且难以诊断和减少,因目标函数通常缺乏直接监督损失用于超参数调优、模型选择与早停。本文引入等序贝尔曼校准,一种一维、模型无关的后处理方法,在保持初始占用率估计排序信息的同时降低此类残差。该方法通过在非递减变换的一维类上进行拟合占用率评估(FORE)来校正估计的尺度与形状。我们证明贝尔曼校准等价于对校准后比率的每种测试函数实现占用平衡的条件固定点性质。进一步推导出校准-精炼界,表明任意拟合比率若校准误差小,则表现接近基于其拟合值的最佳后处理。针对等序贝尔曼校准,建立了有限样本校准保证及相对于初始估计最优单调变换的KLOracle不等式。因此,该方法在下游目标占用函数(包括策略价值估计)上实现小校准误差与小KL风险,误差在最佳单调修正的统计可接受范围内。
原文摘要 · Abstract (English)
Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equation. Existing minimax, primal-dual, and fitted fixed-point estimators can leave residual occupancy-balance violations because of function-class approximation, regularization, or incomplete optimization. These violations are difficult to diagnose and reduce because the objectives generally lack a direct supervised validation loss for hyperparameter tuning, model selection, and early stopping. We introduce isotonic Bellman calibration, a one-dimensional, model-agnostic post-processing method that reduces these violations while preserving the ranking information in any initial occupancy-ratio estimate. The method corrects the estimate's scale and shape by applying fitted occupancy-ratio evaluation (FORE) over a one-dimensional class of nondecreasing transformations. We characterize Bellman calibration as a conditional fixed-point property equivalent to occupancy-balance against every test function of the calibrated ratio. More generally, we derive a calibration-refinement bound showing that any fitted ratio with small calibration error performs nearly as well as the best post-processing based on its fitted values. For isotonic Bellman calibration, we establish finite-sample calibration guarantees and a KL oracle inequality relative to the best monotone transformation of the initial estimate. Consequently, isotonic Bellman calibration achieves small calibration error and KL risk within statistical error of the best monotone correction, with guarantees for downstream target-occupancy functionals, including policy-value estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。