arXiv:2503.06026cs.ROcs.AI2025-03被引 2

用视觉语言模型实现零样本插销插入,无需训练即可识别孔位并估计姿态。

Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models

  • 基于视觉语言模型识别未见过的插销与孔位匹配关系。
  • 在多种新组合上达到90.2%孔位识别准确率,真实场景插入成功率达88.3%。
  • 适用于工业连接器、玩具拼装等复杂场景,适合需要强泛化的机器人装配任务。

实现零样本插销插入——在未见过的插销-孔对上完成插入而无需特定任务训练——仍是机器人领域的基础挑战。该任务要求感知系统具备高度泛化能力,能检测潜在孔位、从多个候选中选出正确匹配、精确估计其姿态,并在不确定性下执行插入。尽管已有学习方法应用于插销插入,但通常无法超越训练时遇到的特定配对。近期视觉语言模型(VLMs)的发展提供了新路径,通过大规模数据实现跨任务鲁棒泛化。受此启发,我们提出一种新型零样本插销插入框架,利用VLM在无几何先验情况下识别匹配孔位并估计其SE(2)姿态。大量实验表明,本方法在多种未见的插销-孔对上(包括3D打印件、玩具拼图和工业连接器)实现了90.2%的孔位识别准确率,显著优于基线。此外,在电脑背板真实连接器插入任务中,系统成功完成孔位检测、正确孔识别、姿态估计与插入,成功率达88.3%。结果凸显了VLM驱动的零样本推理在实现稳健、泛化性强的机器人装配中的潜力。

原文摘要 · Abstract (English)

Achieving zero-shot peg insertion, where inserting an arbitrary peg into an unseen hole without task-specific training, remains a fundamental challenge in robotics. This task demands a highly generalizable perception system capable of detecting potential holes, selecting the correct mating hole from multiple candidates, estimating its precise pose, and executing insertion despite uncertainties. While learning-based methods have been applied to peg insertion, they often fail to generalize beyond the specific peg-hole pairs encountered during training. Recent advancements in Vision-Language Models (VLMs) offer a promising alternative, leveraging large-scale datasets to enable robust generalization across diverse tasks. Inspired by their success, we introduce a novel zero-shot peg insertion framework that utilizes a VLM to identify mating holes and estimate their poses without prior knowledge of their geometry. Extensive experiments demonstrate that our method achieves 90.2% accuracy, significantly outperforming baselines in identifying the correct mating hole across a wide range of previously unseen peg-hole pairs, including 3D-printed objects, toy puzzles, and industrial connectors. Furthermore, we validate the effectiveness of our approach in a real-world connector insertion task on a backpanel of a PC, where our system successfully detects holes, identifies the correct mating hole, estimates its pose, and completes the insertion with a success rate of 88.3%. These results highlight the potential of VLM-driven zero-shot reasoning for enabling robust and generalizable robotic assembly.

机器人零样本视觉语言模型装配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。