用残差学习让机器人更精准地模仿人类动作,完成抓取和行走。
ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
- 先用通用动作数据生成基础动作,再用残差网络精细优化。
- 在仿真和真实机器人上成功率显著提升,训练更快更稳定。
- 适合做机器人抓取、搬运等复杂任务的研发人员参考。
类人机器人全身心运动操控有望变革日常服务与仓储任务。尽管通用运动追踪(GMT)技术已能复现多样人类动作,但现有策略缺乏执行操作任务所需的精度与物体感知能力。为此,我们提出ResMimic,一种两阶段残差学习框架,从人类运动数据中实现精确且富有表现力的类人控制。首先,基于大规模纯人类动作数据训练的GMT策略作为无任务依赖的基础,生成类人全身运动;随后,通过高效而精确的残差策略对GMT输出进行优化,以提升运动性能并融入物体交互。为促进高效训练,我们设计了:(i) 基于点云的物体追踪奖励,实现更平滑优化;(ii) 接触奖励,增强机器人与物体间的准确互动;(iii) 基于课程学习的虚拟物体控制器,稳定早期训练过程。我们在仿真环境和真实Unitree G1机器人上评估了ResMimic,结果表明其在任务成功率、训练效率和鲁棒性方面均显著优于强基线模型。视频展示见https://resmimic.github.io/。
原文摘要 · Abstract (English)
Humanoid whole-body loco-manipulation promises transformative capabilities for daily service and warehouse tasks. While recent advances in general motion tracking (GMT) have enabled humanoids to reproduce diverse human motions, these policies lack the precision and object awareness required for loco-manipulation. To this end, we introduce ResMimic, a two-stage residual learning framework for precise and expressive humanoid control from human motion data. First, a GMT policy, trained on large-scale human-only motion, serves as a task-agnostic base for generating human-like whole-body movements. An efficient but precise residual policy is then learned to refine the GMT outputs to improve locomotion and incorporate object interaction. To further facilitate efficient training, we design (i) a point-cloud-based object tracking reward for smoother optimization, (ii) a contact reward that encourages accurate humanoid body-object interactions, and (iii) a curriculum-based virtual object controller to stabilize early training. We evaluate ResMimic in both simulation and on a real Unitree G1 humanoid. Results show substantial gains in task success, training efficiency, and robustness over strong baselines. Videos are available at https://resmimic.github.io/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。