用算子理论解决连续时间离线强化学习的误差问题
Operator Models for Continuous-Time Offline Reinforcement Learning
- 基于控制扩散过程的无穷小生成器构建世界模型
- 证明价值函数全局收敛并给出有限样本误差界
- 适合研究连续控制与离线学习的学者参考
连续时间随机过程广泛存在于自然与工程系统中。在医疗、自动驾驶和工业控制领域,直接与环境交互往往不安全或不可行,因此需从历史数据中进行离线强化学习。然而,现有方法对离线数据学习策略所固有的近似误差缺乏统计理解。本文将强化学习与哈密顿-雅可比-贝尔曼方程关联,提出一种基于再生核希尔伯特空间中控制扩散过程无穷小生成器的算子理论算法。通过结合统计学习与算子理论,我们建立了价值函数的全局收敛性,并推导出与系统平滑性、稳定性等性质相关的有限样本误差界。理论与数值结果表明,算子方法在基于连续时间最优控制的离线强化学习中具有潜力。
原文摘要 · Abstract (English)
Continuous-time stochastic processes underlie many natural and engineered systems. In healthcare, autonomous driving, and industrial control, direct interaction with the environment is often unsafe or impractical, motivating offline reinforcement learning from historical data. However, there is limited statistical understanding of the approximation errors inherent in learning policies from offline datasets. We address this by linking reinforcement learning to the Hamilton-Jacobi-Bellman equation and proposing an operator-theoretic algorithm based on a simple dynamic programming recursion. Specifically, we represent our world model in terms of the infinitesimal generator of controlled diffusion processes learned in a reproducing kernel Hilbert space. By integrating statistical learning methods and operator theory, we establish global convergence of the value function and derive finite-sample guarantees with bounds tied to system properties such as smoothness and stability. Our theoretical and numerical results indicate that operator-based approaches may hold promise in solving offline reinforcement learning using continuous-time optimal control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。