arXiv:2509.26633cs.ROcs.AI2025-09被引 117

通过保留人-物-环境交互关系,生成高保真机器人动作数据。

OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction

论文配图:OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
图 1 · 摘自论文原文
  • 基于交互网格建模人机物空间关系,保持接触与空间约束。
  • 生成超8小时轨迹,显著减少足部滑移和穿透现象。
  • 支持跨机器人、地形、物体配置的高效数据增强,适合复杂任务训练。

教导类人机器人掌握复杂技能的主流方法是将人类运动重定向为运动学参考,用于训练强化学习(RL)策略。然而,现有重定向流程常因人机身体差异导致物理上不合理的伪影,如足部滑移和穿透。更重要的是,常见方法忽略了对表达性移动和运动操作至关重要的丰富人-物及人-环境交互。为此,我们提出OmniRetarget,一种基于交互网格的数据生成引擎,显式建模并保留代理、地形与操作物体之间的关键空间和接触关系。通过最小化人体与机器人网格间的拉普拉斯形变,并施加运动学约束,OmniRetarget生成运动学可行的轨迹。此外,保留任务相关交互可实现高效数据增强,从单一示范生成适用于不同机器人本体、地形和物体配置的轨迹。我们通过在OMOMO、LAFAN1及自建动作捕捉数据集上重定向运动,生成超过8小时的轨迹,其运动学约束满足度与接触保持能力优于广泛使用的基线方法。此类高质量数据使本体感知强化学习策略仅用5个奖励项与简单的领域随机化,在Unitree G1类人机器人上成功执行长达30秒的攀爬与运动操作技能,无需学习课程。

原文摘要 · Abstract (English)

A dominant paradigm for teaching humanoid robots complex skills is to retarget human motions as kinematic references to train reinforcement learning (RL) policies. However, existing retargeting pipelines often struggle with the significant embodiment gap between humans and robots, producing physically implausible artifacts like foot-skating and penetration. More importantly, common retargeting methods neglect the rich human-object and human-environment interactions essential for expressive locomotion and loco-manipulation. To address this, we introduce OmniRetarget, an interaction-preserving data generation engine based on an interaction mesh that explicitly models and preserves the crucial spatial and contact relationships between an agent, the terrain, and manipulated objects. By minimizing the Laplacian deformation between the human and robot meshes while enforcing kinematic constraints, OmniRetarget generates kinematically feasible trajectories. Moreover, preserving task-relevant interactions enables efficient data augmentation, from a single demonstration to different robot embodiments, terrains, and object configurations. We comprehensively evaluate OmniRetarget by retargeting motions from OMOMO, LAFAN1, and our in-house MoCap datasets, generating over 8-hour trajectories that achieve better kinematic constraint satisfaction and contact preservation than widely used baselines. Such high-quality data enables proprioceptive RL policies to successfully execute long-horizon (up to 30 seconds) parkour and loco-manipulation skills on a Unitree G1 humanoid, trained with only 5 reward terms and simple domain randomization shared by all tasks, without any learning curriculum.

机器人控制动作重定向强化学习交互建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。