无需训练即可识别变形纸盒并自动选点抓取,适合机器人分拣场景。
Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface Scoring

- 用视觉语言模型+几何评分分离识别与抓取点选择
- 单物体抓取成功率88.2%,复杂场景下整体拾取率达72.6%
- 不依赖训练数据,适用于形变大、杂乱环境的纸盒分拣
由于目标物体具有可变形性和几何不一致性,机器人对可回收垃圾的分拣极具挑战。本文提出一种无需训练的吸盘抓取系统,用于分拣变形的无菌饮料纸盒,将目标识别与抓取点选择解耦。通过开放词汇视觉语言模型根据文本提示检测纸盒,SAM2将每个检测结果细化为实例掩码,再利用几何评分方法结合表面平坦度与法向对齐性选取最优吸盘点。对比了三种几何方法:k近邻PCA、Sobel叉积和RANSAC平面拟合。在真实机器人上针对三种形变程度和35个杂乱场景进行评估,单物体抓取成功率达到88.2%,端到端拾取成功率72.6%。
原文摘要 · Abstract (English)
Robotic sorting of recyclable waste is challenging due to the deformable and geometrically inconsistent nature of target objects. We present a training-free suction grasping system for sorting deformed aseptic beverage cartons, decoupling target identification from grasp-point selection. An open-vocabulary vision-language model detects cartons from a text prompt, SAM2 refines each detection into an instance mask, and a geometric scoring method selects the suction point by combining surface flatness with normal alignment. Three geometric methods are compared: k-nearest-neighbour PCA, Sobel cross-product, and RANSAC plane fitting. Evaluated on a real robot across three deformation levels and 35 cluttered scenes, single-object grasp success reaches 88.2% and end-to-end retrieval in clutter is 72.6%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。