arXiv:2608.10332eess.SYcs.LG2026-08

为可微预测控制提供确定性安全保证,无需在线过滤器。

Topological Feasibility Guarantees for Differentiable Predictive Control

论文配图:Topological Feasibility Guarantees for Differentiable Predictive Control
图 1 · 摘自论文原文
  • 基于拓扑分析构建安全可达集,实现离线策略的确定性可行性保障。
  • 训练样本越多,约束违反率单调下降至零,实证验证理论结果。
  • 适合关注学习型控制安全性的研究者和工业控制系统设计者。

可微预测控制(DPC)是一种自监督学习方法,用于近似显式模型预测控制(MPC)策略,相较于基于在线优化的MPC具有显著计算优势。然而,当前可行性保障要么是概率性的,要么依赖在线安全过滤器。本文通过新型拓扑分析方法,对诱导的安全可达集进行研究,首次在不使用在线安全过滤器的情况下,为离线策略优化提供了确定性可行性保证。利用DPC中可微系统动态直接嵌入计算图的模型驱动特性,从拓扑与几何角度分析了学习到的控制策略及其对应系统状态的性质。受此理论启发,提出一种新的自监督离线策略学习方法,采用带有控制屏障函数(CBFs)的代理损失。关键在于,这些性质不仅显著提升策略训练效果,还使从有限训练样本中推导出严格的确定性可行性保证成为可能。大量闭环仿真验证了理论结论,显示经验约束违规随训练样本数量增加而单调递减至零。最终表明,DPC策略优化能生成传统黑箱方法(如强化学习或基于监督学习的近似MPC)无法获得的正式安全证书,为学习型控制中的可行性保障提供了新视角。

原文摘要 · Abstract (English)

Differentiable predictive control (DPC), a self-supervised learning approach for approximating explicit model predictive control (MPC) policies, offers significant computational advantages over online optimization-based MPC. However, feasibility guarantees, a core requirement for safe control, are currently provided either probabilistically or via online safety filters. The lack of rigorous feasibility guarantees for offline policy optimization remains an open problem. This paper establishes deterministic feasibility guarantees for DPC using a novel topological analysis of the induced reachable safe set, without requiring online safety filters. By exploiting the inherent model-based nature of DPC, in which differentiable system dynamics are embedded directly into the computational graph, we analyze the properties of the learned control policies and the corresponding system states from topological and geometric perspectives. Inspired by our theoretical analysis, we propose a novel self-supervised offline policy learning strategy that utilizes a proxy loss with Control Barrier Functions (CBFs). Crucially, these properties not only significantly improve policy training but also enable the derivation of strict, deterministic feasibility guarantees from a finite number of training samples. Extensive closed-loop simulations validate our theoretical findings, demonstrating that the empirical constraint violations monotonically decrease to zero as the training sample size increases. Ultimately, this work illustrates that DPC policy optimization yields formal safety certificates that are structurally unattainable with conventional black-box methods, e.g., reinforcement learning (RL) or supervised learning-based approximate MPC, thereby providing a new perspective on feasibility guarantees in learning-based control.

控制理论安全控制可微控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。