让高维系统安全地学习反馈控制策略
End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers

- 用算子分裂+无雅可比反向传播实现端到端训练
- 在状态维数达1200、控制维数400的系统中保持安全约束
- 适合需要严格安全保证的高维多智能体控制场景
我们研究在控制屏障函数(CBFs)强制的安全约束下,学习高维半全局反馈控制器的问题。将CBF融入端到端策略训练需嵌入基于二次规划的安全滤波器作为优化层,但计算与反向传播瓶颈使以往方法仅限于低维系统(通常不超过16个状态维度)。本文通过结合算子分裂与近期提出的无雅可比反向传播(JFB)方法,实现可扩展的端到端训练,同时通过CBF安全滤波器保留硬性安全保证。理论分析采用非光滑分析技术验证该方法,实证展示其在状态维度高达1200、控制维度达400的非线性多智能体控制问题中的有效性。
原文摘要 · Abstract (English)
We consider the problem of learning high-dimensional semi-global feedback controllers under hard safety constraints enforced by control barrier functions (CBFs). Incorporating CBFs into end-to-end policy training requires embedding a quadratic-program-based safety filter as an optimization layer, but computational and differentiation bottlenecks have largely restricted prior approaches to low-dimensional systems, typically with at most 16 state dimensions. We address this limitation by combining operator splitting with the recently developed Jacobian-Free Backpropagation (JFB) method to enable scalable end-to-end training while preserving hard safety guarantees through the CBF safety filter. We justify this training methodology theoretically using nonsmooth analysis techniques and demonstrate its effectiveness on high-dimensional multi-agent nonlinear control problems with state and control dimensions up to 1200 and 400, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。