无需贝尔曼完备性,直接估计离线强化学习中的状态转移比例。
Fitted Occupancy-Ratio Evaluation without Bellman Completeness
- 通过邻接贝尔曼递归构建拟合固定点方法,直接求解密度比目标。
- 仅需可实现的折扣占用率,即可保证在相对熵下收敛。
- 适用于覆盖不全场景,提供保守的策略价值下界,适合安全评估。
占用率可纠正离线强化学习中的分布偏移,是离线策略评估的核心。现有原始-对偶和极小极大方法通常通过在评价函数类上施加占用平衡矩来估计这些比率。本文提出拟合占用率评估(FORE),一种基于拟合固定点的方法,通过邻接贝尔曼递归刻画折扣占用率。每轮迭代中,FORE在单步转移数据上求解单一层次的密度比目标,将邻接贝尔曼像投影到对数比率类中,以最小化KL散度。与拟合Q值评估不同,其核心近似条件仅为折扣占用率本身的可实现性,无需值函数可实现性或贝尔曼完备性。在该条件下,总体KL投影递归在相对熵意义下收缩至真实比率,因邻接贝尔曼算子为KL收缩映射。对于经验递归,我们建立了有限样本后悔界,表明其在达到近似误差范围内收敛,统计误差由比率假设类的复杂度决定。当全覆盖失效时,引入覆盖停止的FORE,聚焦于首次未覆盖状态动作对前累积的折扣占用率,对非负奖励情形提供目标策略价值的保守下界。拟合比率支持通过奖励重加权、占用加权拟合Q评估以及结合拟合Q函数的双重稳健估计进行直接值估计。结果表明,折扣占用率可实现性即足以保证离线策略评估,无需任何完备性假设。
原文摘要 · Abstract (English)
Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class. We propose fitted occupancy-ratio evaluation (FORE), a fitted fixed-point method that characterizes the discounted occupancy ratio through an adjoint Bellman recursion. At each iteration, FORE solves a single-level density-ratio objective on one-step-transition data, thereby projecting the adjoint Bellman image onto a log-ratio class in Kullback-Leibler (KL) divergence. Unlike analyses of fitted Q-evaluation, which typically require value-function realizability together with Bellman completeness or projected-operator stability, our central approximation condition is just realizability of the discounted occupancy ratio itself. Under this condition, the population KL-projected recursion contracts in relative entropy toward the true ratio by virtue of the adjoint Bellman operator being a KL-contraction. For the empirical recursion, we establish finite-sample regret bounds that yield convergence in KL up to approximation error and a statistical error governed by the complexity of the ratio hypothesis class. When full coverage fails, we introduce coverage-stopped FORE, which targets the discounted occupancy accumulated before the first uncovered state-action pair and yields a conservative lower bound on target-policy value for nonnegative rewards. The fitted ratio supports direct value estimation by reward reweighting, occupancy-weighted fitted Q-evaluation, and doubly robust estimation that combines the fitted ratio with a fitted Q-function. Together, these results identify discounted occupancy-ratio realizability as a sufficient condition for offline policy evaluation without any completeness assumptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。