arXiv:2606.04829cs.RO2026-06被引 1

统一多种运动模态,让机器人一次训练就能模仿不同动作。

M3imic: Learning a Versatile Whole-Body Controller for Multimodal Motion Mimicking

论文配图:M3imic: Learning a Versatile Whole-Body Controller for Multimodal Motion Mimicking
图 1 · 摘自论文原文
  • 用专用编码器将关节角度、人体姿态等不同数据映射到同一空间。
  • 仿真中对未见任务成功率超98%,实现在真实机器人上跨模态迁移。
  • 适合需要多任务通用控制的机器人研发者使用。

构建通用全身控制器对于实现人形机器人在多样化下游任务中的运动能力至关重要,包括行走和人机协同操作。不同任务依赖不同的运动参考模态:行走主要依赖协调的机器人关节轨迹,而操作则需要精确的末端执行器轨迹跟踪。现有方法常忽视密集关节角与稀疏末端位姿之间的表征差异。为此,我们提出多模态模仿(M3imic),一个可泛化的多模态全身控制框架,通过模态专用编码器将机器人关节角、人体姿态轨迹和末端执行器位姿等异构运动参考模态统一映射至共享隐空间。在仿真中采用大规模强化学习训练单一策略,实现跨多种运动参考模态的仿真到现实迁移,无需针对每种模态重新训练。在Unitree G1机器人上开展的大量仿真与真实世界实验表明,该策略在未见测试集上的峰值成功率达98.42%,展现出卓越的泛化能力。代码已开源:https://github.com/Renforce-Dynamics/MultiModalWBC。

原文摘要 · Abstract (English)

Building a general-purpose whole-body controller is essential for enabling diverse motion capabilities in humanoid robots across a wide range of downstream tasks, including locomotion and loco-manipulation. Different tasks rely on distinct motion reference modalities: locomotion primarily depends on coordinated robot joint trajectories, whereas manipulation requires precise end-effector trajectory tracking. Existing methods often overlook the representational mismatch between dense robot joint angles and sparse end-effector poses. To address this, we propose Multi-Modal Mimic (M3imic), a versatile multi-modal whole-body control framework that unifies heterogeneous motion reference modalities, including robot joint angles, human pose trajectories, and end-effector poses, using modality-specific encoders to map them into a shared latent space. Leveraging large-scale reinforcement learning in the simulator, we train a single policy that achieves sim-to-real transfer across multiple motion reference modalities without modality-specific retraining. Extensive simulation and real-world experiments on the Unitree G1 robot are conducted to evaluate the proposed framework. In simulation, the policy achieves a peak success rate of 98.42\% on an unseen test dataset, demonstrating its exceptional generalization capability. The code is available at https://github.com/Renforce-Dynamics/MultiModalWBC

全身控制多模态仿人机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。