用视觉-本体融合建模手指与物体的对应关系,提升机器人抓握灵巧度。
CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World
- 通过6D姿态和本体感知构建交互点云,建立手物对应关系
- 在6个真实任务中达顶尖表现,显著超越现有方法
- 适合需要高精度灵巧操作的机器人研发人员
实现机器人接近人类水平的灵巧操作是该领域的重要目标。近年来基于3D的模仿学习取得进展,但高质量3D表征面临两大挑战:(1) 单视角相机捕获的点云受分辨率、位置及手部遮挡影响大;(2) 全局点云缺乏关键接触信息和空间对应关系,难以支持精细操作。为此,我们提出CordViP框架,利用物体稳健的6D姿态估计和机器人本体感知,构建并学习对应关系。首先引入交互感知点云,建立物体与手部间的对应;随后用于预训练策略,结合以物体为中心的接触图和手臂协调信息,有效捕捉时空动态。实验表明,该方法在6个真实世界任务中达到当前最优性能,显著优于其他基线。结果还显示其对不同物体、视角和场景具有优异泛化性与鲁棒性。代码与视频见https://aureleopku.github.io/CordViP。
原文摘要 · Abstract (English)
Achieving human-level dexterity in robots is a key objective in the field of robotic manipulation. Recent advancements in 3D-based imitation learning have shown promising results, providing an effective pathway to achieve this goal. However, obtaining high-quality 3D representations presents two key problems: (1) the quality of point clouds captured by a single-view camera is significantly affected by factors such as camera resolution, positioning, and occlusions caused by the dexterous hand; (2) the global point clouds lack crucial contact information and spatial correspondences, which are necessary for fine-grained dexterous manipulation tasks. To eliminate these limitations, we propose CordViP, a novel framework that constructs and learns correspondences by leveraging the robust 6D pose estimation of objects and robot proprioception. Specifically, we first introduce the interaction-aware point clouds, which establish correspondences between the object and the hand. These point clouds are then used for our pre-training policy, where we also incorporate object-centric contact maps and hand-arm coordination information, effectively capturing both spatial and temporal dynamics. Our method demonstrates exceptional dexterous manipulation capabilities, achieving state-of-the-art performance in six real-world tasks, surpassing other baselines by a large margin. Experimental results also highlight the superior generalization and robustness of CordViP to different objects, viewpoints, and scenarios. Code and videos are available on https://aureleopku.github.io/CordViP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。