arXiv:2502.18842cs.ROcs.AI2025-02被引 4

用注意力机制融合CLIP与SAM,提升便利店商品精准分割。

Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation

  • 通过梯度注意力引导CLIP与SAM协同工作
  • 在自建数据集上实现高精度物体掩码生成
  • 适合需要精准抓取的机器人操作场景

本文提出一种新型流水线,用于提升机器人在便利店商品货架场景中物体分割的精度。该方法整合了CLIP与SAM两大先进模型,聚焦于多模态数据(图像与文本)的协同利用。通过梯度注意力机制与定制化数据集微调,显著优化了分割性能。尽管CLIP、SAM和Grad-CAM均为成熟组件,但其在本结构化流程中的有效集成具有重要创新性。最终生成的分割掩码可直接作为机器人系统的输入,支持更精准、自适应的物体操作,在便利店场景中具备实用价值。

原文摘要 · Abstract (English)

This paper introduces a novel pipeline to enhance the precision of object masking for robotic manipulation within the specific domain of masking products in convenience stores. The approach integrates two advanced AI models, CLIP and SAM, focusing on their synergistic combination and the effective use of multimodal data (image and text). Emphasis is placed on utilizing gradient-based attention mechanisms and customized datasets to fine-tune performance. While CLIP, SAM, and Grad- CAM are established components, their integration within this structured pipeline represents a significant contribution to the field. The resulting segmented masks, generated through this combined approach, can be effectively utilized as inputs for robotic systems, enabling more precise and adaptive object manipulation in the context of convenience store products.

物体分割机器人操作多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。