arXiv:2512.23805stat.MLcs.LG2025-12

通过加权方法提升离线策略评估稳定性,无需贝尔曼完备性假设。

Fitted Q-Evaluation without Bellman Completeness via Occupancy Weighting

  • 用目标策略的折扣占有比率加权回归,修正投影范数
  • 在有限样本下实现收敛,且对函数类误设和比率估计误差更鲁棒
  • 适合研究离线强化学习评估的学者,尤其关注理论保障者

拟合Q值评估(FQE)是标准的基于回归的离线策略评估方法,但在分布偏移下,仅靠值函数可表示性不足以保证收敛,现有分析常需贝尔曼完备性。我们发现该不稳定性源于几何错配:标准FQE在离线分布诱导的范数下投影贝尔曼目标,未必保持贝尔曼收缩性。为此,我们提出占据加权FQE,仅改变回归权重。以目标策略的折扣占有比率加权,使投影范数与目标策略动态一致,恢复了总体投影贝尔曼算子的收缩性。我们推导了使用估计占据比率时的有限样本保证,分离了有限迭代、统计、近似和比率估计误差。精确占据加权可消除贝尔曼完备性需求;使用估计权重时,近似完备性和值函数可表示性降低对比率估计误差的敏感性,精确可表示性带来更高阶依赖。将占据加权FQE与拟合占据比率评估结合,获得端到端保证,由值函数和占据比率类的复杂度及直接近似误差决定。在覆盖条件下,两类的联合可表示性足以实现一致性估计,无需贝尔曼或批评者侧完备性。受控实验展示了投影范数机制及收缩与覆盖之间的有限样本权衡。

原文摘要 · Abstract (English)

Fitted \(Q\)-evaluation (FQE) is a standard regression-based method for off-policy evaluation, but under distribution shift, value-function realizability alone does not ensure convergence, and existing analyses often require Bellman completeness. We trace this instability to a geometric mismatch: standard FQE projects Bellman targets in the norm induced by the offline distribution, which need not preserve Bellman contraction. We therefore study \emph{occupancy-weighted FQE}, which changes only the regression weights. Weighting by a target-policy discounted occupancy ratio aligns the projection norm with the target-policy dynamics and restores contraction of the population projected Bellman operator. We derive finite-sample guarantees with estimated occupancy ratios and function-class misspecification, separating finite-iteration, statistical, approximation, and ratio-estimation errors. Exact occupancy weighting removes the need for Bellman completeness; with estimated weights, approximate completeness and value-function realizability reduce sensitivity to ratio-estimation error, with exact realizability yielding higher-order dependence. Combining occupancy-weighted FQE with fitted occupancy-ratio evaluation gives an end-to-end guarantee governed by the complexities and direct approximation errors of the value-function and occupancy-ratio classes. Under coverage, joint realizability of these two classes suffices for consistent estimation without Bellman or critic-side completeness. Controlled experiments illustrate the projection-norm mechanism and the finite-sample tradeoff between contraction and coverage.

强化学习离线评估理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。