用神经网络求解连续时间HJB方程,提升强化学习稳定性与精度。
Stabilized neural Hamilton--Jacobi--Bellman solvers: Error analysis and applications in model-based reinforcement learning

- 将有限差分法与神经网络结合,通过随机采样优化残差。
- 在64维线性系统等任务中,误差随模型偏差和策略不匹配上升但可控。
- 适合需要高精度反馈控制的复杂系统强化学习研究者使用。
物理信息神经求解器为连续时间建模强化学习提供新路径,其中最优反馈控制由哈密顿-雅可比-贝尔曼(HJB)方程决定。实际实现常处于既非经典网格法也非连续偏微分方程物理信息神经网络(PINN)的混合区域:价值函数由神经网络表示,有限差分HJB策略评估算子通过网络在偏移点上的查询计算,残差通过随机连续配点最小化。该方法保留了稳定有限差分策略评估结构,同时避免网格值未知量。本文建立了该混合范式下的误差理论。将有限差分视为作用于神经网络的平移算子,证明了一步策略评估在学习动态下的总体$L^2$稳定性估计。该界分离出残差误差、初值与外部边界层不匹配、策略不匹配及模型识别误差,并显式给出学习动态的梯度放大因子,而底层线性评估稳定性无隐藏逆粘性爆炸问题。进一步给出有限样本配点验证与贪心策略改进下的多步传播条件结果。在64维紧凑控制线性二次调节器(LQR)、Allen-Cahn控制、摆杆、霍珀机器人及3D四旋翼基准测试中,对比代表性基于模型与无模型强化学习基线,验证了预测的残差、策略不匹配与学习模型误差趋势。
原文摘要 · Abstract (English)
Physics-informed neural solvers offer a promising route to model-based reinforcement learning in continuous time, where optimal feedback synthesis is governed by Hamilton--Jacobi--Bellman (HJB) equations. Practical implementations often occupy a regime that is neither a classical grid method nor a continuous-PDE PINN: the value function is represented by a neural network, finite-difference HJB policy-evaluation operators are evaluated by network queries at shifted points, and residuals are minimized by random continuous collocation. This regime preserves the stabilized finite-difference policy-evaluation structure while avoiding grid-based value unknowns. We develop an error theory for this hybrid regime. Interpreting finite differences as shift operators acting on neural networks, we prove a population $L^2$ stability estimate for one policy-evaluation step with learned dynamics. The bound separates residual error, initial and exterior-collar mismatch, policy mismatch, and model-identification error, with an explicit gradient amplification factor for learned dynamics, while the underlying linear evaluation stability remains free of hidden inverse-viscosity blow-up. We further give a finite-sample collocation certificate and a conditional multi-step propagation result through greedy policy improvement. Experiments on compact-control LQR upto 64 dimensions, Allen--Cahn control, pendulum, Hopper, and 3D quadrotor benchmarks compare against representative model-based and model-free RL baselines, demonstrating the predicted residual, policy-mismatch, and learned-model error trends.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。