用解耦方法把人类视频转成机器人能用的动作,让机器人学会人类动作。
Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing

- 将人类动作视频分解为任务与身体特征两个独立潜在空间。
- 生成的机器人动作视频在时间和形态上都保持一致,无需成对数据。
- 适合想用互联网人类视频训练机器人的研究者和工程师。
从人类视频学习机器人操作是解决机器人数据瓶颈的有前景方案,但人与机器人之间的分布差异仍是关键挑战。现有方法常产生纠缠表征,使任务信息与人类特有运动学耦合,限制适应性。本文提出一种生成式跨体感视频编辑框架,直接通过显式解耦任务与体感表征来应对该问题。方法通过双对比目标将示范视频分解为两个正交潜在空间:最小化两空间间互信息以确保独立性,同时最大化空间内一致性以生成稳定表征。一个参数高效的适配器将这些潜在码注入冻结的视频扩散模型,仅需单个真人示范即可合成连贯的机器人执行视频,且无需成对跨体感数据。实验表明,该方法生成的机器人示范视频在时间上连贯、形态上准确,为利用网络规模的人类视频实现机器人学习提供可扩展解决方案。
原文摘要 · Abstract (English)
Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical challenge. Existing approaches often produce entangled representations, where task-relevant information is coupled with human-specific kinematics, limiting their adaptability. We propose a generative framework for cross-embodiment video editing that directly addresses this by learning explicitly disentangled task and embodiment representations. Our method factorizes a demonstration video into two orthogonal latent spaces by enforcing a dual contrastive objective: it minimizes mutual information between the spaces to ensure independence while maximizing intra-space consistency to create stable representations. A parameter-efficient adapter injects these latent codes into a frozen video diffusion model, enabling the synthesis of a coherent robot execution video from a single human demonstration, without requiring paired cross-embodiment data. Experiments show our approach generates temporally consistent and morphologically accurate robot demonstrations, offering a scalable solution to leverage internet-scale human video for robot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。