让机器人更真实模仿人类动作,还能应对意外环境。
HybridMimic: Hybrid RL-Centroidal Control for Humanoid Motion Mimicking
- 用学习策略动态控制物理模型,实时预测接触状态和质心速度。
- 在真实机器人上使基底位置误差降低13%,比顶尖强化学习方法更稳。
- 适合需要可靠动作执行的复杂人形机器人任务。
运动模仿通过引导控制策略模仿人类动作,有助于人形机器人在强化学习下完成复杂任务。尽管标准强化学习框架表现出色,但在部署时常忽略机器人动力学,导致在分布外环境中产生不切实际的指令。结合模型基础原理的混合方法可提升性能,但现有方法通常依赖预设接触时机,灵活性不足。本文提出 HybridMimic 框架,其中学习到的策略通过预测连续接触状态和期望质心速度,动态调节基于质心模型的控制器。该架构利用质心动力学的物理基础,生成即使在领域偏移下也可行的前馈力矩。通过引入物理感知奖励,策略被训练为高效使用质心控制器的优化能力,输出精确控制目标与参考力矩。在 Booster T1 人形机器人上的硬件实验表明,相比最先进的强化学习基线,平均基底位置跟踪误差减少 13%,验证了动力学感知部署的鲁棒性。
原文摘要 · Abstract (English)
Motion mimicking, i.e., encouraging the control policy to mimic human motion, facilitates the learning of complex tasks via reinforcement learning (RL) for humanoid robots. Although standard RL frameworks demonstrate impressive locomotion agility, they often bypass explicit reasoning about robot dynamics during deployment, which is a design choice that can lead to physically infeasible commands when the robot encounters out-of-distribution environments. By integrating model-based principles, hybrid approaches can improve performance; however, existing methods typically rely on predefined contact timing, limiting their versatility. This paper introduces HybridMimic, a framework in which a learned policy dynamically modulates a centroidal-model-based controller by predicting continuous contact states and desired centroidal velocities. This architecture exploits the physical grounding of centroidal dynamics to generate feedforward torques that remain feasible even under domain shift. Using physics-informed rewards, the policy is trained to efficiently utilize the centroidal controller's optimization by outputting precise control targets and reference torques. Through hardware experiments on the Booster T1 humanoid, HybridMimic reduces the average base position tracking error by 13\% compared to a state-of-the-art RL baseline, demonstrating the robustness of dynamics-aware deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。