arXiv:2502.18104cs.CV2025-02被引 8

用文本提示构建跨模态特征,提升光学与雷达图像匹配的泛化能力。

PromptMID: Modal Invariant Descriptors Based on Diffusion and Vision Foundation Models for Optical-SAR Image Matching

  • 基于土地利用提示,融合扩散模型与视觉基础模型生成跨域特征。
  • 在四个区域数据集上超越现有方法,新域表现更优。
  • 适合遥感图像匹配、跨模态学习研究者参考。

理想的图像匹配应实现未见域下的稳定高效性能。然而,现有基于学习的光学-SAR图像匹配方法虽在特定场景有效,但泛化能力有限,难以适应实际应用。反复训练或微调模型以应对域差异不仅不够优雅,还带来额外计算开销和数据成本。近年来,通用基础模型展现出提升泛化潜力,但自然图像与遥感图像间视觉域差异限制其直接应用。因此,如何有效利用基础模型增强光学-SAR图像匹配的泛化性仍是挑战。为此,我们提出PromptMID,一种基于土地利用分类先验信息的文本提示方法,构建光学与SAR图像匹配的模态不变描述符。PromptMID通过预训练扩散模型与视觉基础模型(VFMs)提取多尺度模态不变特征,并设计特殊特征聚合模块,有效融合不同粒度特征。在来自四个不同区域的光学-SAR图像数据集上的大量实验表明,PromptMID优于现有最优方法,在已见与未见域均取得优异结果,表现出强跨域泛化能力。源代码将公开于 https://github.com/HanNieWHU/PromptMID。

原文摘要 · Abstract (English)

The ideal goal of image matching is to achieve stable and efficient performance in unseen domains. However, many existing learning-based optical-SAR image matching methods, despite their effectiveness in specific scenarios, exhibit limited generalization and struggle to adapt to practical applications. Repeatedly training or fine-tuning matching models to address domain differences is not only not elegant enough but also introduces additional computational overhead and data production costs. In recent years, general foundation models have shown great potential for enhancing generalization. However, the disparity in visual domains between natural and remote sensing images poses challenges for their direct application. Therefore, effectively leveraging foundation models to improve the generalization of optical-SAR image matching remains challenge. To address the above challenges, we propose PromptMID, a novel approach that constructs modality-invariant descriptors using text prompts based on land use classification as priors information for optical and SAR image matching. PromptMID extracts multi-scale modality-invariant features by leveraging pre-trained diffusion models and visual foundation models (VFMs), while specially designed feature aggregation modules effectively fuse features across different granularities. Extensive experiments on optical-SAR image datasets from four diverse regions demonstrate that PromptMID outperforms state-of-the-art matching methods, achieving superior results in both seen and unseen domains and exhibiting strong cross-domain generalization capabilities. The source code will be made publicly available https://github.com/HanNieWHU/PromptMID.

图像匹配跨模态基础模型遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。