用物体为中心的3D运动场,从人类视频中学习机器人控制。
Object-centric 3D Motion Field for Robot Learning from Human Videos
- 以物体为中心的3D运动场表示动作,提升视频理解精度。
- 3D运动估计误差降低50%以上,任务成功率超55%。
- 适合零样本迁移,能学习精细操作如插入动作。
从人类视频中学习机器人控制是扩展机器人学习的重要方向。然而,如何从视频中提取动作知识(或动作表示)仍是关键挑战。现有方法如视频帧、像素光流和点云光流存在建模复杂或信息丢失等局限。本文提出使用物体为中心的3D运动场来表示动作,并构建新框架从视频中提取该表示以实现零样本控制。提出两个创新组件:一是训练一个去噪3D运动场估计器,能从含噪声深度的人类视频中鲁棒地提取精细物体3D运动;二是密集物体中心3D运动场预测架构,兼顾跨具身迁移与背景泛化能力。在真实场景中评估,实验表明,本方法相比最新方法将3D运动估计误差降低超过50%,在多样任务中平均成功率达55%,而先前方法失败率超90%(≤10%),并可习得插入等细粒度操作技能。
原文摘要 · Abstract (English)
Learning robot control policies from human videos is a promising direction for scaling up robot learning. However, how to extract action knowledge (or action representations) from videos for policy learning remains a key challenge. Existing action representations such as video frames, pixelflow, and pointcloud flow have inherent limitations such as modeling complexity or loss of information. In this paper, we propose to use object-centric 3D motion field to represent actions for robot learning from human videos, and present a novel framework for extracting this representation from videos for zero-shot control. We introduce two novel components in its implementation. First, a novel training pipeline for training a ''denoising'' 3D motion field estimator to extract fine object 3D motions from human videos with noisy depth robustly. Second, a dense object-centric 3D motion field prediction architecture that favors both cross-embodiment transfer and policy generalization to background. We evaluate the system in real world setups. Experiments show that our method reduces 3D motion estimation error by over 50% compared to the latest method, achieve 55% average success rate in diverse tasks where prior approaches fail~($\lesssim 10$\%), and can even acquire fine-grained manipulation skills like insertion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。