无需训练,用概念向量实现零样本物体位姿估计
ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- 利用视觉语言模型生成开放词汇3D概念图,点位带概念向量
- 零样本下相对位姿估计平均ADD(-S)得分提升62%领先基线
- 适合无标注数据、快速部署的机器人与视觉任务
物体位姿估计是计算机视觉与机器人学的基础任务,但多数方法需大量特定数据集训练。与此同时,大规模视觉语言模型展现出强大的零样本能力。本文提出ConceptPose框架,实现完全无需训练、无需模型的零样本物体位姿估计。该方法利用视觉语言模型构建开放词汇3D概念图,每个点由显著性图生成的概念向量标记。通过建立鲁棒的3D-3D对应关系,实现精确的6自由度相对位姿估计。在常见零样本相对位姿估计基准上,本方法未进行任何对象或数据集特定训练,即达到当前最优性能,平均ADD(-S)得分较最强基线提升62%,优于采用大量特定数据训练的方法。
原文摘要 · Abstract (English)
Object pose estimation is a fundamental task in computer vision and robotics, yet most methods require extensive, dataset-specific training. Concurrently, large-scale vision language models show remarkable zero-shot capabilities. In this work, we bridge these two worlds by introducing ConceptPose, a framework for object pose estimation that is both training-free and model-free. ConceptPose leverages a vision-language-model (VLM) to create open-vocabulary 3D concept maps, where each point is tagged with a concept vector derived from saliency maps. By establishing robust 3D-3D correspondences across concept maps, our approach allows precise estimation of 6DoF relative pose. Without any object or dataset-specific training, our approach achieves state-of-the-art results on common zero shot relative pose estimation benchmarks, outperforming the strongest baseline by a relative 62\% in average ADD(-S) score, including methods that utilize extensive dataset-specific training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。