无需训练,单视角重建透明物体3D形状,提升抓取成功率
SR3D: Unleashing Single-view 3D Reconstruction for Transparent and Specular Object Grasping
- 用单张图像生成3D网格,结合视图与关键点匹配定位物体
- 在仿真和真实场景中实现高精度深度重建,支持有效抓取
- 无需训练,适合实际部署的机器人抓取任务
近年来,3D机器人操作在日常物体抓取方面取得进展,但透明和镜面材质仍因深度感知限制而难以处理。尽管已有多种3D重建与深度补全方法应对这些挑战,但普遍存在设置复杂或观测信息利用不足的问题。为此,我们提出一种无需训练的框架SR3D,仅需单视角观测即可实现对透明与镜面物体的机器人抓取。给定单视角RGB与深度图像,SR3D首先通过外部视觉模型基于RGB图像生成3D物体网格;关键思路是利用2D与3D固有的语义与几何信息,通过视图匹配与关键点匹配机制,精确确定物体在原始深度受损3D场景中的位姿与尺度,从而恢复准确的3D深度图,实现有效抓取检测。仿真与真实世界实验均验证了SR3D的重建有效性。
原文摘要 · Abstract (English)
Recent advancements in 3D robotic manipulation have improved grasping of everyday objects, but transparent and specular materials remain challenging due to depth sensing limitations. While several 3D reconstruction and depth completion approaches address these challenges, they suffer from setup complexity or limited observation information utilization. To address this, leveraging the power of single view 3D object reconstruction approaches, we propose a training free framework SR3D that enables robotic grasping of transparent and specular objects from a single view observation. Specifically, given single view RGB and depth images, SR3D first uses the external visual models to generate 3D reconstructed object mesh based on RGB image. Then, the key idea is to determine the 3D object's pose and scale to accurately localize the reconstructed object back into its original depth corrupted 3D scene. Therefore, we propose view matching and keypoint matching mechanisms,which leverage both the 2D and 3D's inherent semantic and geometric information in the observation to determine the object's 3D state within the scene, thereby reconstructing an accurate 3D depth map for effective grasp detection. Experiments in both simulation and real world show the reconstruction effectiveness of SR3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。