arXiv:2506.06535cs.RO2025-06被引 1

用掩码引导特征池化,让机器人更高效地听懂指令抓取新物体。

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

  • 通过掩码引导池化视觉语言特征,减少计算量提升效率。
  • 在OCID-VLG上比之前方法高7%准确率,真实场景成功率73%。
  • 开源大尺寸数据集RefGraspNet,适合开放词汇抓取研究者使用。

通过自然语言指令操控未知物体的机器人操作仍具挑战性。语言驱动机器人抓取(LDRG)从自然语言查询和RGB-D图像中预测稳定抓取位姿。本文提出MapleGrasp框架,利用掩码引导特征池化实现高效视觉-语言驱动抓取。两阶段训练首先基于CLIP视觉-语言特征预测分割掩码,第二阶段在掩码内聚合特征生成像素级抓取预测,提升效率并降低计算开销。引入掩码池化使在OCID-VLG基准上相比先前方法提升7%。此外,我们构建了规模达现有数据集八倍的开源数据集RefGraspNet,显著增强模型对开放词汇抓取的泛化能力。MapleGrasp在RefGraspNet基准上达到89%抓取准确率。该方法在LIBERO基准上性能接近更大规模视觉-语言-动作模型,且对未见任务泛化能力更强。真实世界实验在Franka机械臂上对未见过物体实现73%成功率,超越竞争基线11个百分点。代码已开源。

原文摘要 · Abstract (English)

Robotic manipulation of unseen objects via natural language commands remains challenging. Language driven robotic grasping (LDRG) predicts stable grasp poses from natural language queries and RGB-D images. We propose MapleGrasp, a novel framework that leverages mask-guided feature pooling for efficient vision-language driven grasping. Our two-stage training first predicts segmentation masks from CLIP-based vision-language features. The second stage pools features within these masks to generate pixel-level grasp predictions, improving efficiency, and reducing computation. Incorporating mask pooling results in a 7% improvement over prior approaches on the OCID-VLG benchmark. Furthermore, we introduce RefGraspNet, an open-source dataset eight times larger than existing alternatives, significantly enhancing model generalization for open-vocabulary grasping. MapleGrasp scores a strong grasping accuracy of 89\% when compared with competing methods in the RefGraspNet benchmark. Our method achieves comparable performance to larger Vision-Language-Action models on the LIBERO benchmark, and shows significantly better generalization to unseen tasks. Real-world experiments on a Franka arm demonstrate 73% success rate with unseen objects, surpassing competitive baselines by 11%. Code is provided in our github repository.

机器人抓取视觉语言高效推理开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。