arXiv:2506.08220cs.CV2025-06NeurIPS被引 6

让图像匹配模型学会泛化到未标注关键点,提升跨图泛化能力。

Jamais Vu: Exposing the Generalization Gap in Supervised Semantic Correspondence

  • 将2D关键点升维至无监督的3D几何空间,构建连续形状流形
  • 在未见关键点上性能超越监督基线,跨数据集泛化更强
  • 适合做物体匹配、少样本几何理解的研究者参考

语义对应(SC)旨在为同一类物体的不同实例建立有意义的匹配。我们指出当前监督式SC方法仍受限于稀疏标注训练关键点,本质上仅是关键点检测器。为此,我们提出一种新方法:通过单目深度估计将2D关键点映射到规范3D空间,构建无需显式3D监督或相机参数的连续规范流形。同时引入SPair-U,扩展SPair-71k并新增关键点标注,以更精准评估泛化能力。实验表明,所提模型在未见关键点上显著优于监督基线,且无监督基线在跨数据集泛化时表现优于监督方法。

原文摘要 · Abstract (English)

Semantic correspondence (SC) aims to establish semantically meaningful matches across different instances of an object category. We illustrate how recent supervised SC methods remain limited in their ability to generalize beyond sparsely annotated training keypoints, effectively acting as keypoint detectors. To address this, we propose a novel approach for learning dense correspondences by lifting 2D keypoints into a canonical 3D space using monocular depth estimation. Our method constructs a continuous canonical manifold that captures object geometry without requiring explicit 3D supervision or camera annotations. Additionally, we introduce SPair-U, an extension of SPair-71k with novel keypoint annotations, to better assess generalization. Experiments not only demonstrate that our model significantly outperforms supervised baselines on unseen keypoints, highlighting its effectiveness in learning robust correspondences, but that unsupervised baselines outperform supervised counterparts when generalized across different datasets.

语义对应泛化能力3D重建少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。