arXiv:2501.11299cs.CV2025-01中稿 · IEEE TIP 2025被引 17

仅用单模态数据训练,让模型在多模态图像匹配中保持稳定表现。

MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching

  • 利用Stable Diffusion预训练特征增强单模态关键点描述子
  • 在5个跨模态数据集上实现零样本泛化,匹配精度提升显著
  • 适合缺乏多模态标注数据的医疗与遥感图像匹配场景

许多关键点检测与描述方法在单模态图像匹配中表现良好,但在多模态数据上常因非线性差异而失效。现有方法通常需对齐的多模态数据来学习模态不变特征,但这类数据获取成本高且不切实际。为此,我们提出MIFNet,一种仅使用单模态训练数据即可学习模态不变特征的网络,用于多模态图像匹配中的关键点描述。通过引入新颖的潜在特征聚合模块和累积混合聚合模块,融合Stable Diffusion模型的预训练特征,提升基础描述子的鲁棒性。在三个眼科视网膜数据集(CF-FA、CF-OCT、EMA-OCTA)和两个遥感数据集(Optical-SAR、Optical-NIR)上的实验表明,MIFNet无需接触目标模态即可学习模态不变特征,具备优异的零样本泛化能力。

原文摘要 · Abstract (English)

Many keypoint detection and description methods have been proposed for image matching or registration. While these methods demonstrate promising performance for single-modality image matching, they often struggle with multimodal data because the descriptors trained on single-modality data tend to lack robustness against the non-linear variations present in multimodal data. Extending such methods to multimodal image matching often requires well-aligned multimodal data to learn modality-invariant descriptors. However, acquiring such data is often costly and impractical in many real-world scenarios. To address this challenge, we propose a modality-invariant feature learning network (MIFNet) to compute modality-invariant features for keypoint descriptions in multimodal image matching using only single-modality training data. Specifically, we propose a novel latent feature aggregation module and a cumulative hybrid aggregation module to enhance the base keypoint descriptors trained on single-modality data by leveraging pre-trained features from Stable Diffusion models. %, our approach generates robust and invariant features across diverse and unknown modalities. We validate our method with recent keypoint detection and description methods in three multimodal retinal image datasets (CF-FA, CF-OCT, EMA-OCTA) and two remote sensing datasets (Optical-SAR and Optical-NIR). Extensive experiments demonstrate that the proposed MIFNet is able to learn modality-invariant feature for multimodal image matching without accessing the targeted modality and has good zero-shot generalization ability. The code will be released at https://github.com/lyp-deeplearning/MIFNet.

多模态匹配关键点描述零样本学习医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。