arXiv:2410.08901cs.RO2024-10被引 6

无需训练即可精准抓取物体功能部位,靠语义与几何信息融合提升抓取效果。

SegGrasp: Zero-Shot Task-Oriented Grasping via Semantic and Geometric Guided Segmentation

  • 用视觉语言模型粗分割+凸分解几何信息融合优化分割质量。
  • 在真实机器人抓取任务中,抓取和分割性能比基线提升超15%。
  • 适合零样本场景下需要快速适配新任务的机器人系统使用。

任务导向抓取要求机器人根据物体的功能部分进行抓取,这对构建能在动态环境中执行复杂任务的先进机器人系统至关重要。本文提出一种无需训练的框架,结合语义与几何先验实现零样本任务导向抓取。该框架名为SegGrasp,首先利用GLIP等视觉语言模型进行粗略分割;随后通过凸分解提供的详细几何信息,采用名为GeoFusion的融合策略提升分割精度;最后由抓取网络基于优化后的分割结果生成有效抓取位姿。我们在分割基准和真实机器人抓取任务上进行了实验,结果表明,SegGrasp在抓取和分割性能上均超过基线15%以上。

原文摘要 · Abstract (English)

Task-oriented grasping, which involves grasping specific parts of objects based on their functions, is crucial for developing advanced robotic systems capable of performing complex tasks in dynamic environments. In this paper, we propose a training-free framework that incorporates both semantic and geometric priors for zero-shot task-oriented grasp generation. The proposed framework, SegGrasp, first leverages the vision-language models like GLIP for coarse segmentation. It then uses detailed geometric information from convex decomposition to improve segmentation quality through a fusion policy named GeoFusion. An effective grasp pose can be generated by a grasping network with improved segmentation. We conducted the experiments on both segmentation benchmark and real-world robot grasping. The experimental results show that SegGrasp surpasses the baseline by more than 15\% in grasp and segmentation performance.

抓取规划零样本语义分割机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。