通过图像匹配提升跨视角定位精度,解决大视角变化下的对应关系问题。
ViewBridge:Revisiting Cross-View Localization from Image Matching
- 从图像匹配角度重构跨视角定位,引入表面模型保证几何一致性。
- 在极端视角下实现细粒度且可靠的匹配,定位精度显著提升。
- 适用于高精度地图、自动驾驶等需要精确位置识别的场景。
跨视角定位旨在通过将地面视角图像与航空或卫星影像对齐,估计其3自由度姿态。现有方法通常通过直接回归或在共享鸟瞰图(BEV)空间中的特征对齐来实现,虽能完成粗略对齐,但在大视角变化下难以建立细粒度且几何可靠对应关系,限制了定位的准确性和可解释性。为此,我们从图像匹配视角重新审视该任务,提出统一框架以增强匹配与定位性能。具体地,引入表面模型,约束BEV特征投影至物理上合理的区域以确保几何一致性;设计SimRefiner模块,自适应优化相似度分布以提升匹配可靠性。为进一步推动该领域研究,我们构建了首个包含32,509对跨视角图像并标注像素级对应关系的基准数据集CVFM。大量实验表明,所提方法可在极端视角下实现几何一致且细粒度的对应关系,并显著提高跨视角定位的准确性和稳定性。
原文摘要 · Abstract (English)
Cross-view localization aims to estimate the 3-DoF pose of a ground-view image by aligning it with aerial or satellite imagery. Existing methods typically address this task through direct regression or feature alignment in a shared bird's-eye view (BEV) space. Although effective for coarse alignment, these methods fail to establish fine-grained and geometrically reliable correspondences under large viewpoint variations, thereby limiting both the accuracy and interpretability of localization results. Consequently, we revisit cross-view localization from the perspective of image matching and propose a unified framework that enhances both matching and localization. Specifically, we introduce a Surface Model that constrains BEV feature projection to physically valid regions for geometric consistency, and a SimRefiner that adaptively refines similarity distributions to enhance match reliability. To further support research in this area, we present CVFM, the first benchmark with 32,509 cross-view image pairs annotated with pixel-level correspondences. Extensive experiments demonstrate that our approach achieves geometry-consistent and fine-grained correspondences across extreme viewpoints and further improves the accuracy and stability of cross-view localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。