用单步扩散模型实现遥感多模态图像快速精准配准
OSDM-MReg: Multimodal Image Registration based One Step Diffusion Model
- 设计单步扩散模型,直接生成跨模态统一表征
- 在OSdataset上注册误差低于现有方法30%以上
- 适合遥感图像融合与跨模态分析任务
多模态遥感图像配准旨在对不同传感器获取的图像进行对齐,以支持数据融合与分析。然而,当面对如合成孔径雷达(SAR)与光学图像之间显著的非线性辐射差异时,现有方法常难以提取模态不变特征。为此,本文提出OSDM-MReg框架,通过图像到图像翻译弥合模态差异。具体地,引入一种单步无对齐目标引导条件扩散模型(UTGOS-CDM),将源图与目标图映射至统一表征域。不同于传统条件扩散模型需数百次迭代推理,本模型在训练中引入逆向翻译目标,使测试时可一步直接预测译出图像,大幅加速配准过程。之后,设计多模态多尺度配准网络(MM-Reg),采用新型多模态融合策略,融合单模态与译出的多模态图像,在多尺度和多模态下提升对齐鲁棒性与精度。在OSdataset上的大量实验表明,OSDM-MReg在配准精度上优于当前最优方法。
原文摘要 · Abstract (English)
Multimodal remote sensing image registration aligns images from different sensors for data fusion and analysis. However, existing methods often struggle to extract modality-invariant features when faced with large nonlinear radiometric differences, such as those between SAR and optical images. To address these challenges, we propose OSDM-MReg, a novel multimodal image registration framework that bridges the modality gap through image-to-image translation. Specifically, we introduce a one-step unaligned target-guided conditional diffusion model (UTGOS-CDM) to translate source and target images into a unified representation domain. Unlike traditional conditional DDPM that require hundreds of iterative steps for inference, our model incorporates a novel inverse translation objective during training to enable direct prediction of the translated image in a single step at test time, significantly accelerating the registration process. After translation, we design a multimodal multiscale registration network (MM-Reg) that extracts and fuses both unimodal and translated multimodal images using the proposed multimodal fusion strategy, enhancing the robustness and precision of alignment across scales and modalities. Extensive experiments on the OSdataset demonstrate that OSDM-MReg achieves superior registration accuracy compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。