arXiv:2607.20674cs.LGcs.SY2026-07被引 1

让高维系统安全地学习反馈控制策略

End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers

论文配图:End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers
图 1 · 摘自论文原文
  • 用算子分裂+无雅可比反向传播实现端到端训练
  • 在状态维数达1200、控制维数400的系统中保持安全约束
  • 适合需要严格安全保证的高维多智能体控制场景

我们研究在控制屏障函数(CBFs)强制的安全约束下,学习高维半全局反馈控制器的问题。将CBF融入端到端策略训练需嵌入基于二次规划的安全滤波器作为优化层,但计算与反向传播瓶颈使以往方法仅限于低维系统(通常不超过16个状态维度)。本文通过结合算子分裂与近期提出的无雅可比反向传播(JFB)方法,实现可扩展的端到端训练,同时通过CBF安全滤波器保留硬性安全保证。理论分析采用非光滑分析技术验证该方法,实证展示其在状态维度高达1200、控制维度达400的非线性多智能体控制问题中的有效性。

原文摘要 · Abstract (English)

We consider the problem of learning high-dimensional semi-global feedback controllers under hard safety constraints enforced by control barrier functions (CBFs). Incorporating CBFs into end-to-end policy training requires embedding a quadratic-program-based safety filter as an optimization layer, but computational and differentiation bottlenecks have largely restricted prior approaches to low-dimensional systems, typically with at most 16 state dimensions. We address this limitation by combining operator splitting with the recently developed Jacobian-Free Backpropagation (JFB) method to enable scalable end-to-end training while preserving hard safety guarantees through the CBF safety filter. We justify this training methodology theoretically using nonsmooth analysis techniques and demonstrate its effectiveness on high-dimensional multi-agent nonlinear control problems with state and control dimensions up to 1200 and 400, respectively.

安全控制高维系统深度强化学习屏障函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。