用相似物体匹配指导未知物抓取,提升单视角抓取鲁棒性。
A Multi-Level Similarity Approach for Single-View Object Grasping: Matching, Planning, and Fine-Tuning
- 通过多层级相似性匹配,结合语义、几何与尺寸特征识别相似物体。
- 基于相似物体的预存抓取知识,规划并优化目标物体的抓取姿态。
- 提出新点云描述符C-FPFH,提升部分观测下的匹配精度,适合机器人抓取应用。
从单视角抓取未知物体仍是机器人领域的挑战,主要源于观测不完整带来的不确定性。尽管大型模型如GraspNet-1Billion已成基准方案,但其对传感噪声和环境变化敏感,泛化能力受限。本文摒弃传统学习框架,提出以相似性匹配为核心的新思路:利用已知物体的相似性引导未知物体的抓取。方法包含三步:1)通过观测物体的视觉特征,在包含多种物体模型的数据库中进行相似性匹配,找出高相似候选;2)利用候选模型的预存抓取知识,为未知目标规划仿效式抓取;3)通过局部微调优化抓取质量。为应对部分观测与噪声带来的不确定性,提出多层级相似性匹配框架,融合语义、几何与尺寸特征。特别提出新型点云几何描述符C-FPFH,实现部分点云与完整模型间的精确匹配。同时引入大语言模型、半定向边界框及基于平面检测的点云配准方法,显著提升单视图条件下的匹配精度。视频见 https://youtu.be/qQDIELMhQmk。
原文摘要 · Abstract (English)
Grasping unknown objects from a single view has remained a challenging topic in robotics due to the uncertainty of partial observation. Recent advances in large-scale models have led to benchmark solutions such as GraspNet-1Billion. However, such learning-based approaches still face a critical limitation in performance robustness for their sensitivity to sensing noise and environmental changes. To address this bottleneck in achieving highly generalized grasping, we abandon the traditional learning framework and introduce a new perspective: similarity matching, where similar known objects are utilized to guide the grasping of unknown target objects. We newly propose a method that robustly achieves unknown-object grasping from a single viewpoint through three key steps: 1) Leverage the visual features of the observed object to perform similarity matching with an existing database containing various object models, identifying potential candidates with high similarity; 2) Use the candidate models with pre-existing grasping knowledge to plan imitative grasps for the unknown target object; 3) Optimize the grasp quality through a local fine-tuning process. To address the uncertainty caused by partial and noisy observation, we propose a multi-level similarity matching framework that integrates semantic, geometric, and dimensional features for comprehensive evaluation. Especially, we introduce a novel point cloud geometric descriptor, the C-FPFH descriptor, which facilitates accurate similarity assessment between partial point clouds of observed objects and complete point clouds of database models. In addition, we incorporate the use of large language models, introduce the semi-oriented bounding box, and develop a novel point cloud registration approach based on plane detection to enhance matching accuracy under single-view conditions. Videos are available at https://youtu.be/qQDIELMhQmk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。