arXiv:2503.22462cs.CV2025-03CVPR被引 7

用单目深度估计提升跨视角图像语义匹配精度

SemAlign3D: Semantic Correspondence between RGB-Images through Aligning 3D Object-Class Representations

  • 通过单目深度与视觉大模型特征构建3D物体类别表征
  • 在SPair-71k上整体[email protected]达88.9%,提升3.3点
  • 适合需要鲁棒跨视角匹配的计算机视觉研究者

语义对应近年得益于大视觉模型(LVM)的进展,但现有方法难以捕捉语义区域间的全局几何关系,导致极端视图变化下性能不可靠。本文利用单目深度估计,构建3D物体类别表征以增强几何感知。首先,基于稀疏标注的图像对应数据集,结合单目深度与LVM特征构建3D表征;其次,设计可梯度下降优化的对齐能量函数,实现3D表征与输入图像中物体实例的对齐。该方法在挑战性数据集SPair-71k上取得当前最优结果,三类任务提升超过10点,整体[email protected]从85.6%提升至88.9%。代码与资源详见https://dub.sh/semalign3d。

原文摘要 · Abstract (English)

Semantic correspondence made tremendous progress through the recent advancements of large vision models (LVM). While these LVMs have been shown to reliably capture local semantics, the same can currently not be said for capturing global geometric relationships between semantic object regions. This problem leads to unreliable performance for semantic correspondence between images with extreme view variation. In this work, we aim to leverage monocular depth estimates to capture these geometric relationships for more robust and data-efficient semantic correspondence. First, we introduce a simple but effective method to build 3D object-class representations from monocular depth estimates and LVM features using a sparsely annotated image correspondence dataset. Second, we formulate an alignment energy that can be minimized using gradient descent to obtain an alignment between the 3D object-class representation and the object-class instance in the input RGB-image. Our method achieves state-of-the-art matching accuracy in multiple categories on the challenging SPair-71k dataset, increasing the [email protected] score by more than 10 points on three categories and overall by 3.3 points from 85.6% to 88.9%. Additional resources and code are available at https://dub.sh/semalign3d.

语义对应3D表征单目深度视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。