arXiv:2510.07773cs.ROcs.AI2025-10被引 3

用人体动作轨迹实现人到机器人的技能迁移,无需配对数据。

Trajectory Conditioned Cross-embodiment Skill Transfer

  • 将人体动作抽象为无形态差异的光流轨迹作为通用运动信号。
  • 在模拟环境上降低39.6%的FVD和36.6%的KVD,提升16.7%成功率。
  • 适用于厨房任务等真实场景,支持跨本体技能转移,适合机器人学习研究者。

从人类示范视频中学习操作技能是一个前景广阔但极具挑战的问题,主要源于人体与机器人末端执行器之间的显著本体差异。现有方法依赖成对数据集或手工设计奖励函数,限制了可扩展性和泛化能力。本文提出TrajSkill框架,实现基于轨迹条件的跨本体技能迁移,使机器人能直接从人类示范视频中习得操作技能。核心思路是将人类动作表示为稀疏光流轨迹,以消除形态差异的同时保留关键动力学信息,形成与本体无关的运动提示。在视觉和文本输入的共同引导下,TrajSkill联合生成时序一致的机器人操作视频,并将其转化为可执行动作,从而实现跨本体技能迁移。大量实验表明,在模拟数据(MetaWorld)上,相比最先进方法,TrajSkill分别降低39.6%的FVD和36.6%的KVD,跨本体成功率达16.7%的提升。真实机器人在厨房操作任务中的实验进一步验证了该方法的有效性,证明了人到机器人的实用技能迁移能力。

原文摘要 · Abstract (English)

Learning manipulation skills from human demonstration videos presents a promising yet challenging problem, primarily due to the significant embodiment gap between human body and robot manipulators. Existing methods rely on paired datasets or hand-crafted rewards, which limit scalability and generalization. We propose TrajSkill, a framework for Trajectory Conditioned Cross-embodiment Skill Transfer, enabling robots to acquire manipulation skills directly from human demonstration videos. Our key insight is to represent human motions as sparse optical flow trajectories, which serve as embodiment-agnostic motion cues by removing morphological variations while preserving essential dynamics. Conditioned on these trajectories together with visual and textual inputs, TrajSkill jointly synthesizes temporally consistent robot manipulation videos and translates them into executable actions, thereby achieving cross-embodiment skill transfer. Extensive experiments are conducted, and the results on simulation data (MetaWorld) show that TrajSkill reduces FVD by 39.6\% and KVD by 36.6\% compared with the state-of-the-art, and improves cross-embodiment success rate by up to 16.7\%. Real-robot experiments in kitchen manipulation tasks further validate the effectiveness of our approach, demonstrating practical human-to-robot skill transfer across embodiments.

技能迁移视频生成机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。