用海森矩阵增强数据,大幅减少求解最优控制问题的训练样本数。
Hessian-augmented Supervised Learning for Hamilton-Jacobi-Bellman PDEs

- 利用庞特里亚金原理生成带海森矩阵的多初始值数据
- 二阶信息使样本量降低一个数量级,精度显著提升
- 适合高维系统,可直接导出反馈控制律
针对具有非线性控制仿射动力学的确定性最优控制问题,提出一种数据驱动方法以逼近价值函数。通过从多个初始条件求解庞特里亚金最大值原理最优性系统,生成包含价值、梯度和海森矩阵的数据集,其中海森矩阵由最优轨迹上的矩阵里卡蒂方程获得。这些高阶信息被用于稀疏多项式基在双曲交叉索引集上的加权最小二乘回归,梯度与海森矩阵为每个样本增加额外线性方程,显著降低样本复杂度。反馈律可从学习到的价值函数解析恢复。在高维情况下,采用部分海森策略控制数据生成成本。该方法在状态维度递增的问题上验证有效,二阶数据增强明显提升近似精度与闭环性能,相比低阶方法最多减少一个数量级的训练样本。
原文摘要 · Abstract (English)
A data-driven method is developed for approximating value functions in deterministic optimal control problems with nonlinear control-affine dynamics. The Pontryagin Maximum Principle optimality system is solved from multiple initial conditions to generate training data consisting of values, gradients, and Hessians of the value function, where Hessian information is obtained from a matrix Riccati equation along optimal trajectories. These quantities augment a weighted least-squares regression over sparse polynomial bases on hyperbolic cross index sets, with gradients and Hessians contributing additional linear equations per sample and substantially reducing sample complexity compared to value-only regression. Feedback laws are recovered analytically from the learned value function. In high dimensions, a partial Hessian strategy controls the cost of data generation. The approach is validated on problems of increasing state dimension, where second-order data augmentation is shown to improve approximation accuracy and closed-loop performance, with up to an order-of-magnitude reduction in the number of training samples required relative to lower-order methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。