用物理信息机器学习框架,同时优化自主系统安全与性能。
A Physics-Informed Machine Learning Framework for Safe and Optimal Control of Autonomous Systems
- 将安全与性能统一为状态约束的最优控制问题
- 通过新框架高效逼近满足HJB方程的值函数
- 支持高置信度安全验证和性能退化概率界
随着自主系统在日常生活中日益普及,确保高性能与可靠安全至关重要。然而,安全与性能常相互冲突,难以协同优化。基于学习的方法(如受限强化学习)虽性能强,但安全仅作为软约束,缺乏形式化保障;而形式化方法(如哈密顿-雅可比可达性分析、控制屏障函数)虽能提供严格安全保证,却往往忽略性能,导致控制器过于保守。本文将安全与性能的协同优化建模为状态约束的最优控制问题,其中性能目标由代价函数表示,安全要求作为状态约束。我们证明所得值函数满足哈密顿-雅可比-贝尔曼(HJB)方程,并提出一种新型物理信息机器学习框架进行高效近似。此外,引入基于保形预测的验证策略,量化学习误差,恢复高置信度安全值函数,并给出性能退化概率上界。通过多个案例研究,验证了该框架在复杂高维自主系统中实现可扩展的安全高效控制器学习的有效性。
原文摘要 · Abstract (English)
As autonomous systems become more ubiquitous in daily life, ensuring high performance with guaranteed safety is crucial. However, safety and performance could be competing objectives, which makes their co-optimization difficult. Learning-based methods, such as Constrained Reinforcement Learning (CRL), achieve strong performance but lack formal safety guarantees due to safety being enforced as soft constraints, limiting their use in safety-critical settings. Conversely, formal methods such as Hamilton-Jacobi (HJ) Reachability Analysis and Control Barrier Functions (CBFs) provide rigorous safety assurances but often neglect performance, resulting in overly conservative controllers. To bridge this gap, we formulate the co-optimization of safety and performance as a state-constrained optimal control problem, where performance objectives are encoded via a cost function and safety requirements are imposed as state constraints. We demonstrate that the resultant value function satisfies a Hamilton-Jacobi-Bellman (HJB) equation, which we approximate efficiently using a novel physics-informed machine learning framework. In addition, we introduce a conformal prediction-based verification strategy to quantify the learning errors, recovering a high-confidence safety value function, along with a probabilistic error bound on performance degradation. Through several case studies, we demonstrate the efficacy of the proposed framework in enabling scalable learning of safe and performant controllers for complex, high-dimensional autonomous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。