提出SPIDER框架,提升跨域图像匹配的鲁棒性。
SPIDER: Spatial Image CorresponDence Estimator for Robust Calibration
- 融合2D与3D特征,分阶段优化匹配精度。
- 在大视角变化下仍保持高匹配准确率,优于当前最佳方法。
- 适合需要跨场景视觉对齐的机器人与三维重建应用。
可靠的图像对应关系是基于视觉的空间感知基础,用于恢复三维结构和相机位姿。然而,在航拍、室内和室外等不同场景间进行无约束特征匹配仍具挑战,因外观、尺度和视角差异显著。传统特征匹配为二维到二维问题;而近期的三维基础模型利用双视图几何提供了空间一致性特征匹配能力。尽管强大,我们发现这些空间一致匹配常集中于主导平面区域(如墙面或地面),对细微几何结构不敏感,尤其在大视角变化下。为深入理解此权衡,我们首先通过线性探针实验评估多种视觉基础模型的匹配性能。基于此洞察,我们提出SPIDER——一种通用特征匹配框架,包含共享特征提取主干与两个专用网络头,分别从粗到细估计2D与3D对应关系。最后,我们构建了一个聚焦大基线无约束场景的图像匹配评估基准。SPIDER显著优于现有最先进方法,展现出强大的通用图像匹配能力。
原文摘要 · Abstract (English)
Reliable image correspondences form the foundation of vision-based spatial perception, enabling recovery of 3D structure and camera poses. However, unconstrained feature matching across domains such as aerial, indoor, and outdoor scenes remains challenging due to large variations in appearance, scale and viewpoint. Feature matching has been conventionally formulated as a 2D-to-2D problem; however, recent 3D foundation models provides spatial feature matching properties based on two-view geometry. While powerful, we observe that these spatially coherent matches often concentrate on dominant planar regions, e.g., walls or ground surfaces, while being less sensitive to fine-grained geometric details, particularly under large viewpoint changes. To better understand these trade-offs, we first perform linear probe experiments to evaluate the performance of various vision foundation models for image matching. Building on these insights, we introduce SPIDER, a universal feature matching framework that integrates a shared feature extraction backbone with two specialized network heads for estimating both 2D-based and 3D-based correspondences from coarse to fine. Finally, we introduce an image-matching evaluation benchmark that focuses on unconstrained scenarios with large baselines. SPIDER significantly outperforms SoTA methods, demonstrating its strong ability as a universal image-matching method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。