arXiv:2511.16330cs.RO2025-11中稿 · ICRA被引 2

用数学方法保证机器人柔顺控制的稳定与安全,避免试错风险。

Safe and Optimal Variable Impedance Control via Certified Reinforcement Learning

  • 将控制策略探索重构为稳定增益轨迹的数学采样,从源头保障安全。
  • 在仿真和真实机器人上均实现稳定跟踪,即使存在模型误差仍能保持性能。
  • 适合需要高可靠性的协作机器人场景,如医疗辅助或精密装配。

强化学习(RL)通过结合动态运动基元(DMPs)与可变阻抗控制(VIC)实现机器人复杂协作技能的学习,但其无模型范式常因阻抗增益时变性导致不稳定和危险探索。本文提出认证高斯流形采样(C-GMS),一种以轨迹为中心的新型强化学习框架,可同时学习DMP与VIC策略,并通过构造保证李雅普诺夫稳定性与执行器可行性。该方法将策略探索重定义为从数学定义的稳定增益调度流形中采样,确保每次策略执行均稳定且物理可实现,无需奖励惩罚或事后验证。此外,理论证明本方法在存在有界模型误差和部署不确定性时仍能保证有界跟踪误差。我们在仿真中验证了有效性,并在真实机器人上成功部署,为复杂环境中可靠自主交互提供了新路径。

原文摘要 · Abstract (English)

Reinforcement learning (RL) offers a powerful approach for robots to learn complex, collaborative skills by combining Dynamic Movement Primitives (DMPs) for motion and Variable Impedance Control (VIC) for compliant interaction. However, this model-free paradigm often risks instability and unsafe exploration due to the time-varying nature of impedance gains. This work introduces Certified Gaussian Manifold Sampling (C-GMS), a novel trajectory-centric RL framework that learns combined DMP and VIC policies while guaranteeing Lyapunov stability and actuator feasibility by construction. Our approach reframes policy exploration as sampling from a mathematically defined manifold of stable gain schedules. This ensures every policy rollout is guaranteed to be stable and physically realizable, thereby eliminating the need for reward penalties or post-hoc validation. Furthermore, we provide a theoretical guarantee that our approach ensures bounded tracking error even in the presence of bounded model errors and deployment-time uncertainties. We demonstrate the effectiveness of C-GMS in simulation and verify its efficacy on a real robot, paving the way for reliable autonomous interaction in complex environments.

强化学习机器人控制安全学习阻抗控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。