用普通第一视角视频生成机器人可学的精准操作数据
AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning

- 通过补全遮挡物体和手部动作,还原真实操作细节
- 轨迹误差更低,手腕运动更平滑,优于现有方法
- 适合想用人类视频训练机械臂的科研与工程人员
人类第一人称视频是人形机器人策略学习的可扩展监督来源,但现有方法在手物遮挡、动作简化或依赖专用采集设备方面存在瓶颈。我们提出AgenticFocus,一种混合现实合成管道,将普通第一视角人类视频转化为机器人可训练的示范数据,通过恢复被遮挡物体几何结构、重建完整手部运动,并利用相机相对对齐与分层合成技术将其重定向至人形机器人身体。生成的数据集包含聚焦的视觉观测与同步的机器人动作及状态。AgenticFocus在轨迹误差上优于跨体态基线,手腕运动更平滑,SPARC得分分别为-5.18、-5.56和-6.05。
原文摘要 · Abstract (English)
Human egocentric video is a scalable supervision source for humanoid policy learning, but current pipelines struggle with hand-object occlusion, oversimplified motion, or specialized capture hardware. We introduce AgenticFocus, a Mixed Reality synthesis pipeline that converts ordinary first-person-view human videos into robot-trainable demonstrations by restoring occluded object geometry, reconstructing full-hand motion, and retargeting it to a humanoid embodiment through camera-relative alignment and layered compositing. The resulting dataset pairs focused visual observations with synchronized robot actions and states. AgenticFocus achieves lower trajectory error and smoother wrist motion than cross-embodiment baselines, with SPARC scores of -5.18 versus -5.56 and -6.05.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。