arXiv:2609.02134cs.ROcs.GR2026-09

无需人工标注,用点云自动对齐人与机器人的动作,实现跨形态高精度动作迁移。

Unified Motion Retargeting for Humanoids with Learned Point Cloud Correspondence

论文配图:Unified Motion Retargeting for Humanoids with Learned Point Cloud Correspondence
图 1 · 摘自论文原文
  • 通过学习密集点云对应关系,摆脱手工设计关键点的限制。
  • 在多种机器人形态和运动场景下均实现更高保真度的动作还原。
  • 适合需要大规模动作数据生成的机器人训练场景。

人形机器人学习越来越依赖将海量多样的人类动作数据转化为高质量的机器人参考轨迹。然而,由于人与机器人在形态、自由度、关节范围和运动学约束上的显著差异,动作迁移面临挑战。现有方法通常通过手工设计的稀疏关键点或身体部位配对建立人-机器人对应关系,导致迁移质量高度依赖人工语义设计,难以扩展到不同动作源和机器人形态,且仅提供稀疏的姿势引导。本文提出统一动作迁移框架(UMR),无需手动设计人-机器人映射即可学习密集点云对应关系。通过将外部点云作为人与机器人之间的统一接口,UMR解耦了动作迁移对源端骨骼语义和目标端拓扑结构的依赖。学习到的密集对应关系为受约束的点云匹配优化提供了细粒度几何锚点,实现了表面级姿态对齐及交互接触的直接传递。实验表明,UMR可统一处理异构动作源、机器人本体及下游场景(从行走到交互),相比现有最优方法,在动作保真度和合理性方面均有提升。因此,UMR为大规模人类动作参考转化为机器人可用训练数据提供了可扩展的基础。

原文摘要 · Abstract (English)

Humanoid learning increasingly relies on transforming vast and diverse human motion data into high-quality robot reference trajectories. However, retargeting human motion to humanoid robots is challenging due to substantial differences in morphology, degrees of freedom, joint ranges, and kinematic constraints between humans and robots. Existing retargeting methods typically address these differences by defining human-robot correspondence through hand-crafted sparse keypoints or body-part pairs. As a result, retargeting quality depends heavily on manual semantic design, limiting scalability across motion sources and robot morphologies and providing only sparse guidance for reproducing detailed poses and interactions. In this paper, we present Unified Motion Retargeting (UMR), a framework that learns dense point cloud correspondence without requiring manually designed human-robot mappings. By treating exterior point clouds as a unified interface between human motion and humanoid robots, UMR decouples retargeting from source-specific skeletal semantics and robot-specific topology. The learned dense correspondence provides fine-grained geometric anchors for constrained point cloud matching optimization, enabling surface-level pose alignment and direct transfer of interaction contacts. Experiments demonstrate that UMR unifies retargeting across heterogeneous motion sources, robot embodiments, and downstream scenarios ranging from locomotion to interaction, while achieving higher motion fidelity and plausibility than state-of-the-art methods. UMR therefore provides a scalable foundation for transforming large-scale human motion references into robot-ready training data.

动作迁移点云匹配人形机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。