无需深度传感器,仅用两张彩色图就能精准抓取物体。
MG-Grasp: Metric-Scale Geometric 6-DoF Grasping Framework with Sparse RGB Observations
- 用两张彩色图像和相机参数重建真实尺度的三维点云
- 在GraspNet-1Billion上达到当前最优抓取成功率
- 适合无深度相机的机器人抓取场景
单视图RGB-D抓取检测仍是6自由度机器人抓取系统的常见选择,通常依赖深度传感器。尽管近期有研究探索纯RGB的6-DoF抓取方法,但其几何表示不准确,难以支持可靠物理抓取。为此,我们提出MG-Grasp,一种全新的无深度传感器6-DoF抓取框架,可实现高质量物体抓取。该方法利用双视图3D基础模型结合相机内参与外参,从稀疏的RGB图像中重建出真实尺度且多视角一致的稠密点云,并生成稳定的6-DoF抓取姿态。在GraspNet-1Billion数据集及真实场景中的实验表明,MG-Grasp在基于RGB的6-DoF抓取方法中达到当前最优(SOTA)性能。
原文摘要 · Abstract (English)
Single-view RGB-D grasp detection remains a common choice in 6-DoF robotic grasping systems, which typically requires a depth sensor. While RGB-only 6-DoF grasp methods has been studied recently, their inaccurate geometric representation is not directly suitable for physically reliable robotic manipulation, thereby hindering reliable grasp generation. To address these limitations, we propose MG-Grasp, a novel depth-free 6-DoF grasping framework that achieves high-quality object grasping. Leveraging two-view 3D foundation model with camera intrinsic/extrinsic, our method reconstructs metric-scale and multi-view consistent dense point clouds from sparse RGB images and generates stable 6-DoF grasp. Experiments on GraspNet-1Billion dataset and real world demonstrate that MG-Grasp achieves state-of-the-art (SOTA) grasp performance among RGB-based 6-DoF grasping methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。