让仿人机器人在运行中实时满足各种安全约束,保持动作流畅不越界。
Constrained Whole-Body Tracking for Humanoid Robots

- 用运动空间控制与屏障函数结合,实时约束机器人动作
- 实测可避免碰撞、不超关节极限、保持重心稳定
- 方法高效可部署,适合需要安全动作的机器人应用
强化学习已使仿人机器人实现出色的整体身体灵活性,但确保安全并满足训练后指定的约束仍具挑战。为此,我们提出ConstrainedMimic,一种利用整体运动学与动力学的控制框架,可在强化学习跟踪策略中实现实时约束执行。通过融合操作空间控制与控制屏障函数(CBFs)原理,实现对运动参考和底层动力学的任意运行时约束。在模拟单位树G1机器人上,使用学习策略进行全身运动跟踪与遥操作实验,验证了其在避免自身及外部障碍物碰撞、遵守关节极限、维持质心稳定方面的有效性。该方法在保持当前接触模式与跟踪目标一致的前提下,最小化对策略能力的限制。本方法完全可微,可在CPU、GPU和TPU上运行,部署频率高达300-500 Hz。所有软件将在发表后免费开放。
原文摘要 · Abstract (English)
Recent advances in reinforcement learning (RL) have demonstrated impressive whole-body agility for humanoid robots, yet ensuring safety and satisfying constraints -- particularly those specified after training -- remains a challenge. Towards this goal, we present ConstrainedMimic, a control framework that leverages whole-body kinematics and dynamics for real-time constraint enforcement within RL tracking policies. By integrating principles from operational space control and control barrier functions (CBFs), we enable the satisfaction of arbitrary runtime constraints on both the kinematic reference motion and the underlying dynamics. In whole-body motion-tracking and teleoperation experiments on a (simulated) Unitree G1 with a learned policy, we demonstrate collision avoidance (both with the robot body and external obstacles), joint limits, and center of mass stability constraints. By remaining consistent with the current contact mode and tracking objectives, we minimally restrict the capabilities of the policy when constraints are active. Our method is fully differentiable, runs on CPU, GPU, and TPU, and can be deployed at up to 300-500 Hz. All software will be freely available upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。