arXiv:2509.03893cs.CV2025-09ICCV被引 2

用视觉语言模型伪标注功能部件,实现跨类别图像的密集功能对应关系预测。

Weakly-Supervised Learning of Dense Functional Correspondences

  • 利用视觉语言模型为多视角图像生成功能部件伪标签。
  • 结合像素级对比学习,同时提取功能与空间知识,提升对应精度。
  • 适用于跨类别形状匹配任务,如机器人抓取与三维重建。

建立图像对之间的密集对应关系对于形状重建和机器人操作等任务至关重要。在跨类别匹配的挑战性场景中,物体的功能(即物体对其他物体产生的影响)可指导对应关系的建立,因为具备特定功能的物体部分往往在形状和外观上具有相似性。基于此观察,我们定义了密集功能对应关系,并提出一种弱监督学习范式来解决该预测任务。核心思路是利用视觉语言模型为多视角图像伪标注功能部件,再通过像素级对比学习,将功能与空间知识融合到新模型中,以实现密集功能对应。此外,我们构建了合成与真实数据集作为评估基准。实验表明,该方法优于使用现成自监督图像表示和接地视觉语言模型的基线方案。

原文摘要 · Abstract (English)

Establishing dense correspondences across image pairs is essential for tasks such as shape reconstruction and robot manipulation. In the challenging setting of matching across different categories, the function of an object, i.e., the effect that an object can cause on other objects, can guide how correspondences should be established. This is because object parts that enable specific functions often share similarities in shape and appearance. We derive the definition of dense functional correspondence based on this observation and propose a weakly-supervised learning paradigm to tackle the prediction task. The main insight behind our approach is that we can leverage vision-language models to pseudo-label multi-view images to obtain functional parts. We then integrate this with dense contrastive learning from pixel correspondences to distill both functional and spatial knowledge into a new model that can establish dense functional correspondence. Further, we curate synthetic and real evaluation datasets as task benchmarks. Our results demonstrate the advantages of our approach over baseline solutions consisting of off-the-shelf self-supervised image representations and grounded vision language models.

功能对应弱监督视觉语言模型跨类别匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。