将SAM扩展为可引导抓取检测的新模型,无需先验知识也能精准抓取各类物体。
GraspSAM: When Segment Anything Model Meets Grasp Detection
- 基于SAM的提示驱动架构,融合可学习适配器与轻量解码器
- 在Jacquard、Grasp-Anything等数据集上达顶尖性能
- 支持点、框、语言等多种提示,适合真实机器人场景
抓取检测需在不依赖物体先验知识的前提下,灵活应对各种形状物体,并提供直观的用户引导控制。本文提出GraspSAM,作为分割一切模型(SAM)的创新扩展,实现提示驱动且类别无关的抓取检测。相较于以往方法受限于小规模训练数据,GraspSAM利用SAM的大规模训练能力与提示分割特性,高效支持目标物体与类别无关抓取。通过适配器、可学习标记嵌入和轻量修改解码器,只需极少微调即可将物体分割与抓取预测统一于一个框架。在Jacquard、Grasp-Anything和Grasp-Anything++等多个数据集上均达到当前最优表现。大量实验验证了其在不同提示类型(如点、框、语言)下的灵活性,凸显其在真实机器人应用中的鲁棒性与有效性。
原文摘要 · Abstract (English)
Grasp detection requires flexibility to handle objects of various shapes without relying on prior knowledge of the object, while also offering intuitive, user-guided control. This paper introduces GraspSAM, an innovative extension of the Segment Anything Model (SAM), designed for prompt-driven and category-agnostic grasp detection. Unlike previous methods, which are often limited by small-scale training data, GraspSAM leverages the large-scale training and prompt-based segmentation capabilities of SAM to efficiently support both target-object and category-agnostic grasping. By utilizing adapters, learnable token embeddings, and a lightweight modified decoder, GraspSAM requires minimal fine-tuning to integrate object segmentation and grasp prediction into a unified framework. The model achieves state-of-the-art (SOTA) performance across multiple datasets, including Jacquard, Grasp-Anything, and Grasp-Anything++. Extensive experiments demonstrate the flexibility of GraspSAM in handling different types of prompts (such as points, boxes, and language), highlighting its robustness and effectiveness in real-world robotic applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。