arXiv:2510.00491cs.ROcs.AI2025-10被引 2

用3D轨迹作中间表示,让机器人学会模仿人类操作动作。

Traj2Action: A Co-Denoising Framework for Trajectory-Guided Human-to-Robot Skill Transfer

  • 以人体末端运动轨迹为桥梁,统一人类与机器人的动作表达
  • 在真实机器人上提升27%性能,且人数据越多效果越明显
  • 适合需要低成本迁移人类技能的机器人应用

真实世界机器人学习多样操作技能严重受限于昂贵且难扩展的遥操作示范。尽管人类视频可提供可扩展替代方案,但人类与机器人形态差异导致知识迁移困难。为此,我们提出Traj2Action框架,利用操作末端的3D轨迹作为统一中间表征,将嵌入其中的操作知识转移至机器人动作。该策略首先基于人类与机器人数据联合学习粗略轨迹,形成高层运动规划;再在协同去噪框架中生成精确、适配机器人的动作(如姿态与夹爪状态)。研究验证了该框架在架构设计、跨任务泛化与数据效率方面的有效性,并揭示了整合人类手部示范数据时机器人策略学习的关键规律。在Franka机器人上的大量真实实验表明,相比π₀基线,该方法在短时与长时任务中分别提升27%和22.25%性能,且随人类数据增加持续获益。

原文摘要 · Abstract (English)

Learning diverse manipulation skills for real-world robots is severely bottlenecked by the reliance on costly and hard-to-scale teleoperated demonstrations. While human videos offer a scalable alternative, effectively transferring manipulation knowledge is fundamentally hindered by the significant morphological gap between human and robotic embodiments. To address this challenge and facilitate skill transfer from human to robot, we introduce Traj2Action, a novel framework that bridges this embodiment gap by using the 3D trajectory of the operational endpoint as a unified intermediate representation, and then transfers the manipulation knowledge embedded in this trajectory to the robot's actions. Our policy first learns to generate a coarse trajectory, which forms a high-level motion plan by leveraging both human and robot data. This plan then conditions the synthesis of precise, robot-specific actions (e.g., orientation and gripper state) within a co-denoising framework. Our work centers on two core objectives: first, the systematic verification of the Traj2Action framework's effectiveness-spanning architectural design, cross-task generalization, and data efficiency and second, the revelation of key laws that govern robot policy learning during the integration of human hand demonstration data. This research focus enables us to provide a scalable paradigm tailored to address human-to-robot skill transfer across morphological gaps. Extensive real-world experiments on a Franka robot demonstrate that Traj2Action boosts the performance by up to 27% and 22.25% over $π_0$ baseline on short- and long-horizon real-world tasks, and achieves significant gains as human data scales in robot policy learning.

技能迁移轨迹规划机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。