用物理信息神经网络求解非线性系统无限时域最优控制问题
A Physics-Informed Learning Framework to Solve the Infinite-Horizon Optimal Control Problem
- 通过有限时域变体求解稳态HJB方程,确保唯一解
- 随着时域增大,解可一致逼近最优价值函数,仿真验证有效
- 无需稳定控制器先验知识,支持非多项式基函数
本文提出一种基于物理信息神经网络(PINNs)的框架,用于求解非线性系统的无限时域最优控制问题。由于PINNs擅长求解一类偏微分方程(PDE),可用来通过求解关联的稳态哈密顿-雅可比-贝尔曼(HJB)方程来学习价值函数。然而,稳态HJB方程通常存在多个解,直接应用PINNs可能导致逼近非最优解。为此,本文改用具有唯一解的有限时域变体,其解在时域增大时一致逼近最优价值函数。同时给出了验证时域是否足够大的算法,并提供一种计算量小、对近似误差鲁棒的扩展方法。与现有方法不同,该方法不依赖稳定控制器先验知识,可使用非多项式基函数,且无需迭代策略评估。仿真结果验证并澄清了理论结论。
原文摘要 · Abstract (English)
We propose a physics-informed neural networks (PINNs) framework to solve the infinite-horizon optimal control problem of nonlinear systems. In particular, since PINNs are generally able to solve a class of partial differential equations (PDEs), they can be employed to learn the value function of the infinite-horizon optimal control problem via solving the associated steady-state Hamilton-Jacobi-Bellman (HJB) equation. However, an issue here is that the steady-state HJB equation generally yields multiple solutions; hence if PINNs are directly employed to it, they may end up approximating a solution that is different from the optimal value function of the problem. We tackle this by instead applying PINNs to a finite-horizon variant of the steady-state HJB that has a unique solution, and which uniformly approximates the optimal value function as the horizon increases. An algorithm to verify if the chosen horizon is large enough is also given, as well as a method to extend it -- with reduced computations and robustness to approximation errors -- in case it is not. Unlike many existing methods, the proposed technique works well with non-polynomial basis functions, does not require prior knowledge of a stabilizing controller, and does not perform iterative policy evaluations. Simulations are performed, which verify and clarify theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。