arXiv:2608.18948cs.RO2026-08

将人类操作视频转为机器人可用的高保真动作数据,实现低成本泛化训练。

RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience

论文配图:RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience
图 1 · 摘自论文原文
  • 通过跨体态适配与3D手部状态恢复,实现人-机动作一致性转换。
  • 构建1400万帧的跨机器人动作数据集(RoboEdit-14M),覆盖7种机械臂形态。
  • 适合机器人学习、具身智能研究者,可直接用于真实场景控制策略训练。

收集机器人手物交互数据成本高昂且依赖具体机体,而大量人类操作视频仍无法用于机器人训练。我们提出RoboEdit,一个将人类操作视频转化为动作一致、物理合理且3D手部状态对齐的机器人视频的编辑工具套件。为实现可扩展监督,引入RoboEdit-ADC自动流程,从多体态的RGB视频中重建并重定向3D交互。该流程生成了包含174,000对视频(共1400万帧)的RoboEdit-14M大规模数据集,覆盖七种机器人形态、多样场景与交互类型。核心编辑引擎RoboEdit-Trans采用跨体态自适应模块,在保持时间连贯性的同时调整外观与运动;并集成3D机器人状态解码器,恢复每帧手部状态以支持结构化运动监督。实验表明,RoboEdit在编辑质量上达到当前最优,并支持下游机器人控制策略在真实操作任务中的应用。最终,RoboEdit套件释放了未标注人类视频的巨大潜力,为通用机器人学习提供可扩展、高保真的视觉与3D运动监督。

原文摘要 · Abstract (English)

Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transforms human manipulation videos into action-consistent, physically plausible robot videos with aligned 3D hand states. To enable scalable supervision, we introduce RoboEdit-ADC, an automatic pipeline that reconstructs and retargets 3D interactions from RGB videos across embodiments. This pipeline generates RoboEdit-14M, a large-scale dataset of 174K aligned video pairs (14M frames) spanning seven robot embodiments, diverse scenes, and interaction types. The core editing engine, RoboEdit-Trans, employs cross-embodiment adaptation modules to preserve temporal coherence while adapting appearance and motion. It further integrates a 3D Robot-State Decoder to recover per-frame hand states for structured motion supervision. Experiments show that RoboEdit achieves state-of-the-art editing quality and supports downstream robot control policies in real-world manipulation tasks. Ultimately, the RoboEdit suite unlocks the vast potential of unlabeled human videos, providing scalable, high-fidelity visual and 3D motion supervision for generalizable robot learning. Project webpage: https://roboedit.github.io/

机器人学习视频编辑3D重建动作迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。