arXiv:2605.06934cs.LG2026-05

用学习到的稳定屏障提升自适应控制,让机器人更安全地应对未知干扰。

Learned Lyapunov Shielding for Adaptive Control

  • 用可学习的李雅普诺夫函数和神经网络估计未建模动力学
  • 在不求解优化问题的情况下实时保证系统安全,误差降低24%~41%
  • 适合需要高安全性与鲁棒性的工业级机械臂控制场景

我们为欧拉-拉格朗日系统增强了一个槽特-李自适应控制器,引入三个可学习组件:基于楚列斯基参数化的结构化二次李雅普诺夫函数 $V_ψ$,添加有界力矩修正的残差软演员-评论家策略,以及估计未建模动力学的物理信息神经网络。从单仿射约束 $\dot V_ψ + \alpha V_ψ \le 0$ 推导出闭式安全滤波器,无需在线求解二次规划即可将策略输出投影至安全集。理论证明:在控制退化集上满足漂移衰减条件下滤波器全局可行;精确屏蔽下系统指数稳定,鲁棒扩展的裕度取决于PINN近似误差;三时间尺度策略-证书-乘子更新几乎必然收敛至KKT点;证书在紧集上具有概率近似正确(PAC)泛化界。在具有非线性摩擦和变负载的2自由度机械臂上,学习到的证书贡献了主要性能提升:在名义摩擦下跟踪误差下降41%,在激进摩擦下下降24%(训练分布中心)。7自由度弗兰卡-埃米卡熊猫机械臂的可扩展性研究确认全管道工业级干净收敛,明确了超越模型基基准收益的条件,并揭示了学习证书的热启动病态现象,对部署具实际意义。

原文摘要 · Abstract (English)

We augment the Slotine--Li adaptive controller for Euler--Lagrange systems with three learned components: a structured-quadratic Lyapunov function \(V_ψ\) whose positive-definiteness follows from a Cholesky parameterization, a residual Soft Actor--Critic policy that adds bounded torque corrections to the analytic baseline, and a physics-informed neural network that estimates unmodeled dynamics. A closed-form safety filter, derived from the single affine constraint \(\dot V_ψ+ αV_ψ\le 0\), projects every policy output onto the safe set without requiring an online QP solver. We prove: global feasibility of the filter under a drift-decay condition on the control-degeneracy set; exponential stability under exact shielding, with a robust extension whose margin depends on the PINN approximation error; almost-sure convergence of the three-timescale policy--certificate--multiplier updates to a KKT point; and a PAC generalization bound for the certificate over compacts. On a 2-DOF manipulator with nonlinear friction and variable payload, the learned certificate accounts for most of the empirical gain: tracking error drops by 41\% on nominal friction and 24\% on aggressive friction at the centroid of the training distribution. A 7-DOF scalability study on a Franka Emika Panda confirms clean convergence of the full pipeline at industrial scale, identifies the conditions under which gains over exact model-based baselines should and should not be expected, and documents a warm-start pathology of the learned certificate that has practical implications for deployment.

自适应控制安全强化学习机器人李雅普诺夫

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。