提出无需已知动态的时序差分方法,直接求解随机连续系统的贝尔曼方程。
A Temporal Difference Method for Stochastic Continuous Dynamics
- 基于时序差分思想,设计模型无关的连续时间更新机制。
- 理想连续动力学下实现指数收敛,实测优于基于转移核的方法。
- 适合研究随机控制与无模型强化学习融合的学者使用。
对于由常微分方程(ODE)和随机微分方程(SDE)建模的连续系统,贝尔曼最优性原理表现为哈密顿-雅可比-贝尔曼(HJB)方程,这是强化学习(RL)的理论目标。尽管近期强化学习进展成功利用该形式,但现有方法通常假设系统动态已知,因需显式访问动力学方程的系数函数以按HJB方程更新价值函数。本文克服了这一固有限制:提出一种仍以HJB方程为目标的无模型方法,并相应地设计了对应的时序差分算法。我们建立了理想连续时间动力学下的指数收敛性,并在实验中展示了其相对于基于转移核方法的潜在优势。该公式为连接随机控制与无模型强化学习开辟了道路。
原文摘要 · Abstract (English)
For continuous systems modeled by dynamical equations such as ODEs and SDEs, Bellman's Principle of Optimality takes the form of the Hamilton-Jacobi-Bellman (HJB) equation, which provides the theoretical target of reinforcement learning (RL). Although recent advances in RL successfully leverage this formulation, the existing methods typically assume the underlying dynamics are known a priori because they need explicit access to the coefficient functions of dynamical equations to update the value function following the HJB equation. We address this inherent limitation of HJB-based RL; we propose a model-free approach still targeting the HJB equation and propose the corresponding temporal difference method. We establish exponential convergence of the idealized continuous-time dynamics and empirically demonstrate its potential advantages over transition-kernel-based formulations. The proposed formulation paves the way toward bridging stochastic control and model-free reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。