arXiv:2507.06428math.OCcs.LG2025-07被引 5

用神经网络求解高维随机控制方程,理论保证收敛且实测可处理200维问题。

Neural Actor-Critic Methods for Hamilton-Jacobi-Bellman PDEs: Asymptotic Analysis and Numerical Studies

  • critic 精心设计边界条件,actor 通过最小化哈密顿积分更新,结构更高效。
  • 理论证明无限宽网络会收敛到精确解,克服了有限网络易陷局部最优的缺陷。
  • 在200维问题上验证有效,尤其适合非凸哈密顿函数等复杂场景。

我们对一种用于求解高维随机控制理论中哈密顿-雅可比-贝尔曼(HJB)偏微分方程的神经演员-评论家机器学习算法进行了数学分析与数值研究。评论家网络的结构确保边界条件始终精确满足(而非通过损失函数强制),并采用有偏梯度以降低计算开销;演员网络通过最小化全域哈密顿量积分进行训练,其估计值由评论家提供。我们证明,当演员与评论家隐藏单元数趋于无穷时,两者的训练动态在Sobolev型空间中收敛至某一无限维常微分方程(ODE)。进一步地,在哈密顿量具有类凸性假设下,该极限ODE的任意不动点均为原随机控制问题的解,为算法性能提供了重要保障——因为有限宽度网络可能因损失函数非凸而仅收敛至局部极小值。数值实验表明,该算法可在高达200维的问题中准确求解。我们构造了一系列具有已知解析解的递增复杂度随机控制问题,涵盖线性二次调节器到高度复杂的非凸哈密顿方程,借此系统评估了该神经演员-评论家方法在求解HJB方程方面的优劣。

原文摘要 · Abstract (English)

We mathematically analyze and numerically study an actor-critic machine learning algorithm for solving high-dimensional Hamilton-Jacobi-Bellman (HJB) partial differential equations from stochastic control theory. The architecture of the critic (the estimator for the value function) is structured so that the boundary condition is always perfectly satisfied (rather than being included in the training loss) and utilizes a biased gradient which reduces computational cost. The actor (the estimator for the optimal control) is trained by minimizing the integral of the Hamiltonian over the domain, where the Hamiltonian is estimated using the critic. We show that the training dynamics of the actor and critic neural networks converge in a Sobolev-type space to a certain infinite-dimensional ordinary differential equation (ODE) as the number of hidden units in the actor and critic $\rightarrow \infty$. Further, under a convexity-like assumption on the Hamiltonian, we prove that any fixed point of this limit ODE is a solution of the original stochastic control problem. This provides an important guarantee for the algorithm's performance in light of the fact that finite-width neural networks may only converge to a local minimizers (and not optimal solutions) due to the non-convexity of their loss functions. In our numerical studies, we demonstrate that the algorithm can solve stochastic control problems accurately in up to 200 dimensions. In particular, we construct a series of increasingly complex stochastic control problems with known analytic solutions and study the algorithm's numerical performance on them. These problems range from a linear-quadratic regulator equation to highly challenging equations with non-convex Hamiltonians, allowing us to identify and analyze the strengths and limitations of this neural actor-critic method for solving HJB equations.

强化学习偏微分方程神经网络随机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。