arXiv:2606.17385cs.RO2026-06被引 3

用网络视频生成可适配任意机器人的4D手物交互数据,让机器人从看视频学会动作。

EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning

论文配图:EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning
图 1 · 摘自论文原文
  • 构建模块化引擎,自动将任意视角视频转为带物理可靠性的4D手物交互数据
  • 生成包含手部轨迹、6-DoF物体位姿和接触状态的度量级4D表示,减少视觉重建漂移
  • 支持跨形态机器人动作重定向,适用于部分可见或不同视角的视频,适合开放世界学习

互联网视频蕴含海量具身人类操作知识,但将任意RGB视频转化为可执行的机器人训练数据仍是主要瓶颈。现有实验室或工厂采集的数据集规模与多样性有限,制约开放世界机器人学习。我们不构建静态数据集,而是提出EgoInfinity——一个通用的4D手物交互数据引擎,实现网页级数据生成以支持机器人重定向与学习。该引擎模块化设计,集成感知、分割、重建、交互感知优化与重定向,无需人工标注即可自动化解决传统上不可扩展的视频到动作问题。其模块结构能持续受益于各组件的技术进步。通过交叉模块度量校准与交互感知优化,提升物理可靠性,显著降低纯视觉重建中的漂移与接触不一致。我们还提出一种新型运动重定向器,将恢复的3D手部运动编译为多种机器人形态的可执行关节轨迹,实现从任意视角与拍摄尺度(如人体仅部分可见)视频到任何机器人动作的视频到动作重定向。在感知保真度、运动学可行性、接触一致性、跨体态泛化及真实机器人技能获取(如抓取、切割、擦拭、倾倒)等方面验证了EgoInfinity的有效性,为从互联网视频到可执行机器人行为建立了可扩展桥梁。

原文摘要 · Abstract (English)

Internet videos constitute the largest reservoir of embodied human manipulation knowledge, yet converting arbitrary RGB footage into actionable robot training data remains a major bottleneck. Existing lab- or factory-collected datasets are narrow in scale and diversity, limiting open-world robot learning. Instead of proposing a static dataset, we introduce EgoInfinity, a universal 4D hand-object interaction data engine that enables web-scale data generation for robot retargeting and learning. EgoInfinity is a modular engine integrating perception, segmentation, reconstruction, interaction-aware refinement, and retargeting to automate this traditionally unscalable video-to-action problem without human-in-the-loop annotation. Its modular design lets the engine continuously benefit from advances in any incorporated component. With EgoInfinity, in-the-wild human manipulation videos are lifted into agent-agnostic, metric 4D hand-object representations, including hand trajectories, 6-DoF object poses, and contact-relevant states. Rather than naively connecting standalone components, EgoInfinity combines cross-module metric calibration with interaction-aware refinement to improve physical reliability, reducing drift and contact inconsistencies common in pure visual reconstruction. We further propose a novel motion retargeter that compiles the recovered 3D hand motions into executable joint trajectories for diverse robot morphologies, enabling video-to-action retargeting on any robot from arbitrary viewpoints and shot sizes (e.g., the human body is only partially visible). We validate EgoInfinity across perception fidelity, kinematic feasibility, contact consistency, cross-embodiment generalization, and real-robot skill acquisition (e.g., grasping, cutting, wiping, and pouring), demonstrating a scalable bridge from internet videos to executable robot behavior for open-world robot learning.

4D交互视频生成机器人学习动作重定向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。