arXiv:2606.12604cs.RO2026-06被引 5

将第一视角人手视频转为高保真机器人操作数据,无需真实机器人示范。

EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations

论文配图:EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations
图 1 · 摘自论文原文
  • 用视觉和动作映射,把人眼视角视频转成机器人可用的观察与动作序列。
  • 在仿真和真实机器人上验证,实现零样本灵巧操作策略学习。
  • 适合想低成本获取机器人演示数据的研究者和开发者。

灵巧操作受限于大规模机器人示范数据的收集成本。第一视角人类视频提供了可扩展且多样化的操作行为来源,但直接用于机器人学习需跨越两个鸿沟:人类与机器人观测之间的视觉差异,以及人类动作与机器人可执行动作之间的行为差距。我们提出 EgoEngine,一个可扩展的框架,将第一视角人类操作视频转化为高保真机器人数据。给定一段第一视角RGB视频,EgoEngine生成:(i) 以机器人视角替换人类、保留场景上下文与时间对齐的高保真机器人观察视频;(ii) 在可行性约束下与任务对齐的可执行机器人动作轨迹。仿真与真实机器人实验表明,EgoEngine实现了人类视频到机器人数据的可扩展转换,并首次在无真实机器人示范的情况下,实现零样本视觉-运动灵巧操作策略学习。项目网站:https://egoengine.github.io。

原文摘要 · Abstract (English)

Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human videos offer a scalable source of diverse manipulation behaviors, but directly using them for robot learning requires bridging two gaps: the visual gap between human and robot observations, and the action gap between human motion and robot-executable action. We propose EgoEngine, a scalable framework for transforming egocentric human manipulation videos into high-fidelity robot data. Given an egocentric RGB video, EgoEngine produces: (i) a high-fidelity robot observation video replacing human with robot while preserving scene context and temporal alignment, and (ii) a task-aligned, executable robot action trajectory under feasibility constraints. Experiments in simulation and on real robots show that EgoEngine enables scalable conversion of human videos into robot data and, to our knowledge, demonstrates the first zero-shot visuomotor dexterous policy learning from egocentric human videos without real-robot demonstrations. Project website: https://egoengine.github.io.

灵巧操作视觉-运动第一视角零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。