系统梳理视觉与跨模态特征匹配方法,涵盖从传统算法到深度学习的演进。
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey
- 区分不同模态设计专用特征匹配策略,如深度图用几何描述子、点云用稀疏/稠密学习
- 基于Transformer的LoFTR等新方法显著提升跨模态匹配鲁棒性,适应性强于传统SIFT
- 适合研究多模态感知、医学图像配准及视觉-语言对齐的开发者参考
特征匹配是计算机视觉的核心任务,对图像检索、立体匹配、3D重建和SLAM等应用至关重要。本综述全面回顾了基于模态的特征匹配技术,涵盖传统手工设计方法与当代深度学习方法,涉及RGB图像、深度图、3D点云、LiDAR扫描、医学图像及视觉-语言交互等多种模态。传统方法依赖Harris角点、SIFT、ORB等检测器与描述子,在中等模态变化下表现稳健,但面对显著模态差异时性能下降。现代深度学习方法,如无检测器的CNN-based SuperPoint与Transformer-based LoFTR,大幅提升了跨模态鲁棒性与适应性。文中重点介绍模态感知进展:深度图使用几何与深度特定描述子,点云采用稀疏与稠密学习,LiDAR扫描引入注意力增强网络,医学图像则发展出MIND等专用描述子。跨模态应用在医学图像配准与视觉-语言任务中凸显特征匹配向多样化数据交互演进的趋势。
原文摘要 · Abstract (English)
Feature matching is a cornerstone task in computer vision, essential for applications such as image retrieval, stereo matching, 3D reconstruction, and SLAM. This survey comprehensively reviews modality-based feature matching, exploring traditional handcrafted methods and emphasizing contemporary deep learning approaches across various modalities, including RGB images, depth images, 3D point clouds, LiDAR scans, medical images, and vision-language interactions. Traditional methods, leveraging detectors like Harris corners and descriptors such as SIFT and ORB, demonstrate robustness under moderate intra-modality variations but struggle with significant modality gaps. Contemporary deep learning-based methods, exemplified by detector-free strategies like CNN-based SuperPoint and transformer-based LoFTR, substantially improve robustness and adaptability across modalities. We highlight modality-aware advancements, such as geometric and depth-specific descriptors for depth images, sparse and dense learning methods for 3D point clouds, attention-enhanced neural networks for LiDAR scans, and specialized solutions like the MIND descriptor for complex medical image matching. Cross-modal applications, particularly in medical image registration and vision-language tasks, underscore the evolution of feature matching to handle increasingly diverse data interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。