arXiv:2505.11494cs.RO2025-05中稿 · the 2025 IEEE/RSJ …被引 7

用学习到的动态模型实现人形机器人运行时的安全保障

SHIELD: Safety on Humanoids via CBFs In Expectation on Learned Dynamics

  • 基于真实数据训练随机残差动态模型,捕捉系统不确定性和行为
  • 通过概率化控制屏障函数,在不重训的前提下实现安全约束
  • 可在不改动原有控制器的情况下部署,适合需要安全性的机器人应用

机器人学习已生成高效的黑箱控制器,用于人形机器人复杂任务如动态行走。然而,确保动态安全性(即约束满足)仍具挑战。强化学习通过奖励工程启发式嵌入约束,修改约束需重新训练;基于模型的方法如控制屏障函数(CBFs)可实现实时约束定义并提供形式化保证,但依赖精确的动力学模型。本文提出SHIELD,一种分层安全框架:(1) 利用名义控制器在硬件上的实际运行数据,训练一个生成式、随机的动态残差模型,捕捉系统行为与不确定性;(2) 在名义(学习的行走)控制器之上添加安全层,通过随机离散时间CBF公式,在概率意义下强制执行安全约束。结果是无需侵入性修改即可实现风险与性能平衡的随机安全保证。在Unitree G1人形机器人上进行的硬件实验表明,结合未知的强化学习控制器与机载感知,SHIELD可安全穿越室内外多变环境并实现避障。

原文摘要 · Abstract (English)

Robot learning has produced remarkably effective ``black-box'' controllers for complex tasks such as dynamic locomotion on humanoids. Yet ensuring dynamic safety, i.e., constraint satisfaction, remains challenging for such policies. Reinforcement learning (RL) embeds constraints heuristically through reward engineering, and adding or modifying constraints requires retraining. Model-based approaches, like control barrier functions (CBFs), enable runtime constraint specification with formal guarantees but require accurate dynamics models. This paper presents SHIELD, a layered safety framework that bridges this gap by: (1) training a generative, stochastic dynamics residual model using real-world data from hardware rollouts of the nominal controller, capturing system behavior and uncertainties; and (2) adding a safety layer on top of the nominal (learned locomotion) controller that leverages this model via a stochastic discrete-time CBF formulation enforcing safety constraints in probability. The result is a minimally-invasive safety layer that can be added to the existing autonomy stack to give probabilistic guarantees of safety that balance risk and performance. In hardware experiments on an Unitree G1 humanoid, SHIELD enables safe navigation (obstacle avoidance) through varied indoor and outdoor environments using a nominal (unknown) RL controller and onboard perception.

人形机器人安全控制强化学习概率保障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。