深度学习重塑图像匹配,实现端到端高精度三维重建。
Deep Learning Reforms Image Matching: A Survey and Outlook
- 用可学习模块替代传统匹配流程各环节
- 端到端模型在姿态恢复等任务上表现更优
- 适合计算机视觉与SLAM研究者参考
图像匹配通过建立两视图图像间的对应关系,恢复三维结构和相机几何,是计算机视觉的核心技术,广泛应用于视觉定位、三维重建和同步定位与地图构建(SLAM)。传统由“检测器-描述子、特征匹配器、异常值过滤器、几何估计器”组成的流水线在复杂场景下表现不佳。近年来,深度学习显著提升了匹配的鲁棒性和精度。本综述从独特视角系统回顾深度学习如何逐步改造经典图像匹配流程。我们的分类体系与传统流程高度契合:一是用可学习模块替代流程中各个独立步骤,包括可学习检测器-描述子、异常值过滤器和几何估计器;二是将多个步骤融合为端到端可学习模块,涵盖中间层稀疏匹配器、端到端半稠密/稠密匹配器和位姿回归器。我们首先分析两类方法的设计原则、优势与局限,随后在相对位姿恢复、单应性估计和视觉定位任务上对代表性方法进行基准测试。最后讨论开放挑战并展望未来研究方向。通过系统分类与评估深度学习驱动的策略,本综述清晰呈现了图像匹配领域的演进图景,并指明了进一步创新的关键路径。
原文摘要 · Abstract (English)
Image matching, which establishes correspondences between two-view images to recover 3D structure and camera geometry, serves as a cornerstone in computer vision and underpins a wide range of applications, including visual localization, 3D reconstruction, and simultaneous localization and mapping (SLAM). Traditional pipelines composed of ``detector-descriptor, feature matcher, outlier filter, and geometric estimator'' falter in challenging scenarios. Recent deep-learning advances have significantly boosted both robustness and accuracy. This survey adopts a unique perspective by comprehensively reviewing how deep learning has incrementally transformed the classical image matching pipeline. Our taxonomy highly aligns with the traditional pipeline in two key aspects: i) the replacement of individual steps in the traditional pipeline with learnable alternatives, including learnable detector-descriptor, outlier filter, and geometric estimator; and ii) the merging of multiple steps into end-to-end learnable modules, encompassing middle-end sparse matcher, end-to-end semi-dense/dense matcher, and pose regressor. We first examine the design principles, advantages, and limitations of both aspects, and then benchmark representative methods on relative pose recovery, homography estimation, and visual localization tasks. Finally, we discuss open challenges and outline promising directions for future research. By systematically categorizing and evaluating deep learning-driven strategies, this survey offers a clear overview of the evolving image matching landscape and highlights key avenues for further innovation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。