arXiv:2605.05791cs.LG2026-05

为强化学习中的拟合Q迭代提供首个连续空间下的有限样本理论框架。

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration

  • 基于测度论与巴拿赫空间,建立适应数据分布的性能界
  • 证明序列Rademacher复杂度控制贝尔曼回归泛化误差
  • 首次给出连续空间下拟合Q迭代的路径累积在线后悔率

尽管强化学习有望革新复杂非线性机器人系统的控制,但其模型无关的离策略深度强化学习在实践上的启发式成功与理论基础之间仍存在深刻鸿沟,后者主要局限于表格型或可线性化的场景。我们识别出这一差距源于三个传统领域的分离:(i) 测度论马尔可夫决策过程在一般空间上的基础限制分析仅针对精确动态规划,忽略学习过程的所有误差来源;(ii) 确定性误差传播分析通过集中系数处理近似误差,但缺乏估计误差的有限样本分析;(iii) PAC泛化界刻画了简化拓扑结构下的估计误差。本文通过统一理论框架,将拟合Q迭代(FQI)推广至一般的可测博雷尔空间。主要成果是在测度论概率与巴拿赫空间中贝尔曼算子压缩性的结合下,获得一个适应数据分布的有限样本性能界。我们证明,序列Rademacher复杂度控制在策略依赖数据采集下的贝尔曼回归泛化误差。进一步将分析扩展至连续空间,首次提供拟合Q迭代的累积路径在线后悔率保证。这些结果为现代深度强化学习算法的正式分析奠定了必要基础。

原文摘要 · Abstract (English)

While reinforcement learning (RL) promises to revolutionize the control of complex nonlinear robotic systems, a profound gap persists between the heuristic success of model-free off-policy deep RL and the underlying theory, which remains largely confined to tabular or linearizable settings. We identify the cause of this gap as an emergent isolation of three traditions: (i) measure-theoretic MDP foundations on general spaces limit their analysis to exact dynamic programming and ignore all error sources of a learning process; (ii) deterministic error propagation analysis addresses the approximation error via concentrability coefficients without a finite-sample analysis of the estimation error; and (iii) PAC generalization bounds characterize the estimation errors of simplified topologies. We bridge these traditions with a unified theoretical framework for fitted Q-iteration (FQI) on general measurable Borel spaces. Our main result provides a finite-sample, adaptive-data performance bound by chaining measure-theoretic probability with Bellman-operator contraction in Banach spaces. We prove that sequential Rademacher complexity controls Bellman-regression generalization under policy-dependent data collection. We further extend this analysis to provide the first cumulative, pathwise online regret guarantee for FQI in continuous spaces. These results lay the necessary foundations for the formal analysis of many modern deep RL algorithms.

强化学习理论分析泛化误差在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。