从网络教学视频中精准追踪物体6D位姿,助力机器人抓取。
6D Object Pose Tracking in Internet Videos for Robotic Manipulation
- 无先验知识下检索相似3D模型并匹配图像,实现无需对象模板的6D位姿估计。
- 通过轨迹优化生成平滑6D运动序列,在真实与仿真环境中成功控制机械臂。
- 适用于互联网视频和第一视角视频,为具身智能提供可迁移的运动数据。
本文旨在从互联网教学视频中提取被操作物体的时序一致6D位姿轨迹。由于拍摄条件不受控、物体运动细微且动态、且目标物体精确网格未知,现有6D位姿估计方法面临挑战。为此,提出三项贡献:首先,设计一种新方法,无需对象先验即可估计任意物体在输入图像中的6D位姿,流程包括(i)从大规模模型库中检索与目标物体相似的CAD模型,(ii)将检索到的模型与输入图像进行6D对齐,(iii)基于场景建立物体绝对尺度。其次,通过帧间精细跟踪,从互联网视频中提取平滑的6D物体轨迹,并利用轨迹优化将其映射至机械臂配置空间。第三,在YCB-V、HOPE-Video及新构建的手动标注6D轨迹教学视频数据集上进行全面评估与消融实验,显著优于现有RGB 6D位姿估计方法。最后,验证所估计的6D运动可成功应用于7轴机械臂,实现在虚拟仿真与真实环境中的控制;并在EPIC-KITCHENS数据集的自拍视角视频上取得成功,展示其在具身智能中的应用潜力。
原文摘要 · Abstract (English)
We seek to extract a temporally consistent 6D pose trajectory of a manipulated object from an Internet instructional video. This is a challenging set-up for current 6D pose estimation methods due to uncontrolled capturing conditions, subtle but dynamic object motions, and the fact that the exact mesh of the manipulated object is not known. To address these challenges, we present the following contributions. First, we develop a new method that estimates the 6D pose of any object in the input image without prior knowledge of the object itself. The method proceeds by (i) retrieving a CAD model similar to the depicted object from a large-scale model database, (ii) 6D aligning the retrieved CAD model with the input image, and (iii) grounding the absolute scale of the object with respect to the scene. Second, we extract smooth 6D object trajectories from Internet videos by carefully tracking the detected objects across video frames. The extracted object trajectories are then retargeted via trajectory optimization into the configuration space of a robotic manipulator. Third, we thoroughly evaluate and ablate our 6D pose estimation method on YCB-V and HOPE-Video datasets as well as a new dataset of instructional videos manually annotated with approximate 6D object trajectories. We demonstrate significant improvements over existing state-of-the-art RGB 6D pose estimation methods. Finally, we show that the 6D object motion estimated from Internet videos can be transferred to a 7-axis robotic manipulator both in a virtual simulator as well as in a real world set-up. We also successfully apply our method to egocentric videos taken from the EPIC-KITCHENS dataset, demonstrating potential for Embodied AI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。