arXiv:2506.05312cs.CV2025-06ICCV被引 16

用3D感知伪标签提升图像间语义对应精度

Do It Yourself: Learning Semantic Correspondence from Pseudo-Labels

  • 通过3D感知链式生成伪标签,优化预训练特征
  • 在SPair-71k上提升超过4%,优于同类方法
  • 适合需要少标注的跨图像匹配任务

在计算机视觉中,跨图像与物体实例间寻找语义相似点的对应关系是长期挑战。尽管大型预训练视觉模型可作为语义匹配的有效先验,但在对称物体或重复部件上仍存在歧义。本文提出通过3D感知伪标签改进语义对应估计。具体地,训练一个适配器,利用3D感知链式生成的伪标签,通过松弛循环一致性过滤错误标签,并施加3D球面原型映射约束。相比以往工作,该方法减少了对数据集特定标注的需求,在SPair-71k数据集上实现了超过4%的绝对提升,且在类似监督条件下相较其他方法提升超7%。所提方法具有通用性,实验验证其可轻松扩展至其他数据源。

原文摘要 · Abstract (English)

Finding correspondences between semantically similar points across images and object instances is one of the everlasting challenges in computer vision. While large pre-trained vision models have recently been demonstrated as effective priors for semantic matching, they still suffer from ambiguities for symmetric objects or repeated object parts. We propose improving semantic correspondence estimation through 3D-aware pseudo-labeling. Specifically, we train an adapter to refine off-the-shelf features using pseudo-labels obtained via 3D-aware chaining, filtering wrong labels through relaxed cyclic consistency, and 3D spherical prototype mapping constraints. While reducing the need for dataset-specific annotations compared to prior work, we establish a new state-of-the-art on SPair-71k, achieving an absolute gain of over 4% and of over 7% compared to methods with similar supervision requirements. The generality of our proposed approach simplifies the extension of training to other data sources, which we demonstrate in our experiments.

语义对应伪标签3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。