arXiv:2603.19632cs.ROcs.SY2026-03中稿 · RA-L journal被引 2

用可微收缩层让强化学习控制四足机器人更稳定且可验证。

ContractionPPO: Certified Reinforcement Learning via Differentiable Contraction Layers

  • 在PPO中加入状态相关收缩度量层,实现性能与稳定性同步优化。
  • 训练出的收缩度量可证明系统在模拟中具增量指数稳定性。
  • 硬件实验表明对强外部扰动仍有认证稳定的鲁棒控制效果。

在非结构化环境中,足式机器人运动不仅需要高性能控制策略,还需形式化保证以应对扰动。传统控制方法常依赖精心设计的参考轨迹,但在高维、接触密集系统(如四足机器人)中构建困难。相比之下,强化学习(RL)直接学习隐含生成运动的策略,并能利用训练时可用的特权信息(如完整状态和动力学),这些信息在部署时不可见。我们提出ContractionPPO框架,通过在近端策略优化(PPO)中引入状态依赖的收缩度量层,实现足式机器人的可认证稳健规划与控制。该方法使策略在最大化性能的同时,生成可证明模拟闭环系统增量指数稳定的收缩度量。该度量由Lipschitz神经网络参数化,并与策略联合训练,可并行或作为PPO主干的辅助头。尽管收缩度量不用于真实执行,我们推导了最坏情况收缩率的上界,证明其从仿真到真实部署具有泛化能力。四足机器人硬件实验表明,ContractionPPO即使在强外部扰动下也能实现鲁棒且可认证稳定的控制。

原文摘要 · Abstract (English)

Legged locomotion in unstructured environments demands not only high-performance control policies but also formal guarantees to ensure robustness under perturbations. Control methods often require carefully designed reference trajectories, which are challenging to construct in high-dimensional, contact-rich systems such as quadruped robots. In contrast, Reinforcement Learning (RL) directly learns policies that implicitly generate motion, and uniquely benefits from access to privileged information, such as full state and dynamics during training, that is not available at deployment. We present ContractionPPO, a framework for certified robust planning and control of legged robots by augmenting Proximal Policy Optimization (PPO) RL with a state-dependent contraction metric layer. This approach enables the policy to maximize performance while simultaneously producing a contraction metric that certifies incremental exponential stability of the simulated closed-loop system. The metric is parameterized as a Lipschitz neural network and trained jointly with the policy, either in parallel or as an auxiliary head of the PPO backbone. While the contraction metric is not deployed during real-world execution, we derive upper bounds on the worst-case contraction rate and show that these bounds ensure the learned contraction metric generalizes from simulation to real-world deployment. Our hardware experiments on quadruped locomotion demonstrate that ContractionPPO enables robust, certifiably stable control even under strong external perturbations.

强化学习机器人控制可认证性稳定控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。