arXiv:2507.04397cs.CV2025-07被引 1

用状态空间模型提升光学与SAR图像配准精度,解决纹理少、跨模态差异大难题。

Multi-Expert Learning Framework with the State Space Model for Optical and SAR Image Registration

  • 设计多专家框架,通过动态融合多变换特征增强跨模态共享表示
  • 引入Mamba状态空间模型,线性复杂度捕获全局上下文,提升配准准确率
  • 适合遥感图像配准研究者,尤其关注跨模态融合与高效模型设计

光学与合成孔径雷达(SAR)图像配准对多模态图像融合至关重要。然而,现有深度学习方法面临三大挑战:(i) 光学与SAR图像间显著的非线性辐射差异影响共享特征学习;(ii) 图像纹理有限导致判别性特征提取困难;(iii) 卷积神经网络(CNN)局部感受野限制上下文信息建模,而Transformer虽能捕捉长程全局特征但计算开销高。为此,本文提出一种基于状态空间模型的多专家学习框架(ME-SSM)用于光学与SAR图像配准。首先,为提升纹理匮乏场景下的配准性能,构建多专家学习框架,从输入图像的不同变换中提取特征,并采用可学习软路由动态融合,丰富特征表达。其次,引入状态空间模型Mamba,通过多方向交叉扫描策略以线性复杂度高效捕捉全局上下文关系,扩展感受野并避免高计算成本。此外,采用多层级特征聚合(MFA)模块增强多尺度特征融合与交互。大量实验验证了所提方法在光学与SAR图像配准任务上的有效性与优势。

原文摘要 · Abstract (English)

Optical and Synthetic Aperture Radar (SAR) image registration is crucial for multi-modal image fusion and applications. However, several challenges limit the performance of existing deep learning-based methods in cross-modal image registration: (i) significant nonlinear radiometric variations between optical and SAR images affect the shared feature learning and matching; (ii) limited textures in images hinder discriminative feature extraction; (iii) the local receptive field of Convolutional Neural Networks (CNNs) restricts the learning of contextual information, while the Transformer can capture long-range global features but with high computational complexity. To address these issues, this paper proposes a multi-expert learning framework with the State Space Model (ME-SSM) for optical and SAR image registration. Firstly, to improve the registration performance with limited textures, ME-SSM constructs a multi-expert learning framework to capture shared features from multi-modal images. Specifically, it extracts features from various transformations of the input image and employs a learnable soft router to dynamically fuse these features, thereby enriching feature representations and improving registration performance. Secondly, ME-SSM introduces a state space model, Mamba, for feature extraction, which employs a multi-directional cross-scanning strategy to efficiently capture global contextual relationships with linear complexity. ME-SSM can expand the receptive field, enhance image registration accuracy, and avoid incurring high computational costs. Additionally, ME-SSM uses a multi-level feature aggregation (MFA) module to enhance the multi-scale feature fusion and interaction. Extensive experiments have demonstrated the effectiveness and advantages of our proposed ME-SSM on optical and SAR image registration.

图像配准多模态状态空间模型遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。