用RL引导的MPPI实现人形机器人高精度全身控制
RGB: RL Guided Whole-Body MPPI for Humanoid Control

- 用预训练RL策略做采样先验,引导MPPI生成动态可行动作
- 在29自由度机器人上实现平均280Hz稳定控制,精度优于纯RL方法
- 无需重训即可添加新任务目标,适合需要灵活响应的复杂场景
人形机器人在接触丰富的环境中需兼具鲁棒性与精确性的全身控制器。尽管深度强化学习(RL)能实现稳定控制,但其行为高度依赖训练目标与指令接口,难以在不重训的前提下新增反馈目标。本文提出一种基于强化学习引导的全身模型预测路径积分(MPPI)框架,作为预训练RL策略的附加反馈控制器。不将RL策略直接用作最终控制器,而是将其作为采样先验,引导MPPI轨迹向动态可行行为偏移。任务目标通过模块化MPPI代价项定义,MPPI通过在线持续修正RL先验,满足这些目标而无需重训策略。在MuJoCo中对29自由度的Unitree G1人形机器人进行仿真,实现了平均280~300Hz的稳定高速控制。该方法在相同指令接口下,相比纯RL基线提升了任务级精度,有效纠正了直线行走中的系统性漂移,并能跟踪额外的全身参考信号。
原文摘要 · Abstract (English)
Humanoid robots require whole-body controllers that are both robust and precise in contact-rich environments. While deep reinforcement learning (RL) achieves robust stability, its behavior is tightly coupled to the training objective and command interface, making it difficult to add new feedback objectives without retraining. In this study, we propose an RL guided whole-body model predictive path integral (MPPI) framework that acts as an add-on feedback controller on top of a pretrained RL policy. Instead of using RL policy as the final controller, we use it as a sampling prior that biases MPPI rollouts toward dynamically feasible behaviors. Task objectives are specified through modular MPPI cost terms, and MPPI closes the loop by continuously correcting the RL prior online to satisfy these objectives without retraining the policy. Simulations on a 29-DoF Unitree G1 humanoid in MuJoCo demonstrate stable high-rate control (average 280~Hz). The proposed method improves task-level precision over a pure RL baseline under the same command interface. This is achieved by correcting systematic drift during straight walking and tracking additional whole-body reference signals imposed through the cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。