证明了椭圆性让连续时间强化学习的函数逼近不再难于监督学习。
Continuous-time reinforcement learning: ellipticity enables model-free value function approximation
- 利用扩散过程的椭圆性,建立贝尔曼算子的新性质。
- 提出基于最小二乘回归的拟合Q-learning算法,误差可控。
- 适用于无模型连续时间控制,适合研究随机过程强化学习者。
我们研究离散时间观测与动作下对连续时间马尔可夫扩散过程的离策略强化学习。考虑无需对动态结构做不切实际假设的无模型函数逼近算法,直接从数据中学习价值和优势函数。借助扩散过程的椭圆性,我们建立了贝尔曼算子在希尔伯特空间中的正定性和有界性新类。基于这些性质,提出Sobolev-prox拟合Q-learning算法,通过迭代求解最小二乘回归问题来学习价值与优势函数。我们推导出估计误差的奥拉克不等式,受以下四项支配:(i) 函数类的最佳逼近误差,(ii) 局部复杂度,(iii) 指数衰减的优化误差,(iv) 数值离散化误差。这些结果表明,椭圆性是使马尔可夫扩散过程的函数逼近强化学习不比监督学习更难的关键结构性质。
原文摘要 · Abstract (English)
We study off-policy reinforcement learning for controlling continuous-time Markov diffusion processes with discrete-time observations and actions. We consider model-free algorithms with function approximation that learn value and advantage functions directly from data, without unrealistic structural assumptions on the dynamics. Leveraging the ellipticity of the diffusions, we establish a new class of Hilbert-space positive definiteness and boundedness properties for the Bellman operators. Based on these properties, we propose the Sobolev-prox fitted $q$-learning algorithm, which learns value and advantage functions by iteratively solving least-squares regression problems. We derive oracle inequalities for the estimation error, governed by (i) the best approximation error of the function classes, (ii) their localized complexity, (iii) exponentially decaying optimization error, and (iv) numerical discretization error. These results identify ellipticity as a key structural property that renders reinforcement learning with function approximation for Markov diffusions no harder than supervised learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。