arXiv:2505.22000eess.IV2025-05

无需标注数据,通过协同学习实现多模态遥感图像精准配准

Collaborative Learning for Unsupervised Multimodal Remote Sensing Image Registration: Integrating Self-Supervision and MIM-Guided Diffusion-Based Image Translation

  • 设计三模块协同训练框架,融合自监督与扩散模型生成一致图像对
  • 在多个数据集上优于主流无监督方法,部分超越有监督基线
  • 适合遥感、医学图像等领域需跨模态配准的场景

多模态遥感图像在辐射、纹理和结构特征上的显著差异,给精确配准带来挑战。尽管有监督深度学习表现良好,但依赖大规模标注数据,实用性受限。传统无监督方法通常通过最小化特征表示差异来优化配准,但在空间和辐射变化较大时难以稳健捕捉几何偏差,导致收敛不稳定。为此,我们提出无监督多模态图像配准的协同学习框架CoLReg,将无监督配准重构为三个组件的协同训练:(1) 跨模态图像翻译网络MIMGCD,采用可学习的最大索引图(MIM)引导的条件扩散模型生成模态一致图像对;(2) 自监督中间配准网络,利用MIMGCD输出生成的精确位移标签估计几何变换;(3) 用中间网络预测的伪标签训练的简化版跨模态配准网络。三者通过交替优化策略相互增强。该协同机制逐步减少模态差异,提升伪标签质量,最终提高配准精度。在多个数据集上的大量实验表明,CoLReg在性能上达到或超过现有最先进无监督方法,甚至超越若干有监督基线。

原文摘要 · Abstract (English)

The substantial modality-induced variations in radiometric, texture, and structural characteristics pose significant challenges for the accurate registration of multimodal images. While supervised deep learning methods have demonstrated strong performance, they often rely on large-scale annotated datasets, limiting their practical application. Traditional unsupervised methods usually optimize registration by minimizing differences in feature representations, yet often fail to robustly capture geometric discrepancies, particularly under substantial spatial and radiometric variations, thus hindering convergence stability. To address these challenges, we propose a Collaborative Learning framework for Unsupervised Multimodal Image Registration, named CoLReg, which reformulates unsupervised registration learning into a collaborative training paradigm comprising three components: (1) a cross-modal image translation network, MIMGCD, which employs a learnable Maximum Index Map (MIM) guided conditional diffusion model to synthesize modality-consistent image pairs; (2) a self-supervised intermediate registration network which learns to estimate geometric transformations using accurate displacement labels derived from MIMGCD outputs; (3) a distilled cross-modal registration network trained with pseudo-label predicted by the intermediate network. The three networks are jointly optimized through an alternating training strategy wherein each network enhances the performance of the others. This mutual collaboration progressively reduces modality discrepancies, enhances the quality of pseudo-labels, and improves registration accuracy. Extensive experimental results on multiple datasets demonstrate that our ColReg achieves competitive or superior performance compared to state-of-the-art unsupervised approaches and even surpasses several supervised baselines.

图像配准多模态扩散模型自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。