arXiv:2503.04545cs.ROcs.CV2025-03被引 3

用预训练ViT提取语义特征,实现无需训练的通用机器人视觉定位

ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual Servoing

  • 利用预训练ViT提取视觉语义特征,融合传统方法与学习方法优势
  • 在扰动场景下相比经典方法提升31.2%,且收敛速度媲美专用训练模型
  • 无需特定任务或物体训练,可泛化到未见物体的抓取与工业操作

视觉伺服使机器人能精确将末端执行器定位至目标物体。传统方法依赖手工特征,虽通用但易受遮挡和环境变化影响;学习方法虽更鲁棒,却需大量训练。本文提出一种基于预训练视觉变换器的视觉伺服方法,通过其提取语义特征,结合两类范式优势,并具备超越样本范围的泛化能力。在无扰动场景中实现完全收敛,在扰动场景中相较经典图像基视觉伺服提升达31.2%相对性能。即使在收敛速率上也达到学习方法水平,且无需任务或物体特定训练。真实世界测试验证了其在末端定位、工业箱体操作及未见物体抓取中的稳健表现,仅需同类别参考图像即可完成。代码与仿真环境已公开:https://alessandroscherl.github.io/ViT-VS/

原文摘要 · Abstract (English)

Visual servoing enables robots to precisely position their end-effector relative to a target object. While classical methods rely on hand-crafted features and thus are universally applicable without task-specific training, they often struggle with occlusions and environmental variations, whereas learning-based approaches improve robustness but typically require extensive training. We present a visual servoing approach that leverages pretrained vision transformers for semantic feature extraction, combining the advantages of both paradigms while also being able to generalize beyond the provided sample. Our approach achieves full convergence in unperturbed scenarios and surpasses classical image-based visual servoing by up to 31.2\% relative improvement in perturbed scenarios. Even the convergence rates of learning-based methods are matched despite requiring no task- or object-specific training. Real-world evaluations confirm robust performance in end-effector positioning, industrial box manipulation, and grasping of unseen objects using only a reference from the same category. Our code and simulation environment are available at: https://alessandroscherl.github.io/ViT-VS/

视觉伺服ViT机器人泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。