用视觉语言模型指导吸盘抓取,让机器人在复杂环境里精准抓物。
SuctionPrompt: Visual-assisted Robotic Picking with a Suction Cup Using Vision-Language Models and Facile Hardware Design
- 通过视觉语言模型提示+3D检测定位抓取点
- 抓取成功率65.0%,吸盘定位准确率75.4%
- 适合工业场景中快速部署的智能抓取任务
大型语言模型和视觉-语言模型(VLMs)的发展推动了机器人系统在多个领域的应用。然而,如何有效将这些模型融入真实世界中的机器人任务仍是一大挑战。我们开发了一种名为SuctionPrompt的通用机器人系统,结合VLM的提示技术与3D检测,实现多样化动态环境中物品抓取。该方法强调将3D空间信息与自适应动作规划相结合,使机器人能够在新环境中接近并操作物体。验证实验表明,系统在常见物品抓取中实现了65.0%的成功率,吸盘定位准确率达到75.4%。研究证明,即使采用简单的3D处理方式,VLMs在机器人操作任务中依然具有显著有效性。
原文摘要 · Abstract (English)
The development of large language models and vision-language models (VLMs) has resulted in the increasing use of robotic systems in various fields. However, the effective integration of these models into real-world robotic tasks is a key challenge. We developed a versatile robotic system called SuctionPrompt that utilizes prompting techniques of VLMs combined with 3D detections to perform product-picking tasks in diverse and dynamic environments. Our method highlights the importance of integrating 3D spatial information with adaptive action planning to enable robots to approach and manipulate objects in novel environments. In the validation experiments, the system accurately selected suction points 75.4%, and achieved a 65.0% success rate in picking common items. This study highlights the effectiveness of VLMs in robotic manipulation tasks, even with simple 3D processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。