arXiv:2606.03341cs.CV2026-06

用结构化状态空间双性模型提升多模态图像配准的效率与精度

Cross-Modality Feature Fusion Based on Structured State Space Duality for Multimodal Image Registration Network

论文配图:Cross-Modality Feature Fusion Based on Structured State Space Duality for Multimodal Image Registration Network
图 1 · 摘自论文原文
  • 引入SSD模块分三尺度提取多模态特征,增强局部结构表达
  • 提出跨模态交互与多尺度渐进融合机制,实现全局特征有效对齐
  • 在VIS-SAR、VIS-IR、VIS-NIR数据集上优于现有方法,兼顾性能与速度

在多模态图像配准中,共享结构信息提取是核心挑战。相较于Transformer,结构化状态空间双性(SSD)在训练和推理中具备更高效率与更强的全局特征提取能力。受此启发,本文提出新型配准网络RegNetMamba-2,将SSD嵌入粗到精匹配流程,有效提取局部与全局结构特征。首先,在三个不同尺度上应用SSD进行多模态特征提取,并通过特征缩放函数强化前景边缘与结构信息。其次,设计基于SSD的跨模态特征融合模型,包含跨模态交互(CMI)模块与多尺度融合(MSF)模块:CMI模块以交叉形式提取各尺度的跨模态特征,MSF模块采用逐层向上融合策略,整合所有尺度的多模态特征以获得精细特征。最终,从1/8尺度的CMI与1/2尺度的MSF中提取特征,计算像素级匹配概率得分,建立对应关系。大量实验表明,与当前最先进的深度学习算法相比,RegNetMamba-2在VIS-SAR(OSDataset)、VIS-IR(LGHD/RoadSence)及VIS-NIR(RGB-NIR sense)数据集上均实现了优异的配准性能与计算效率。

原文摘要 · Abstract (English)

In multi-modal image registration, the primary challenge lies in shared structural information extraction. Compared to Transformers, Structured State Space Duality (SSD) offers greater global structural feature extraction with higher efficiency during training and inference. Inspired by these advantages, we propose a novel algorithm for multi-modal image registration, named RegNetMamba-2. Our algorithm incorporates SSD into coarse-to-fine matching process to extract local and global structural features effectively. Firstly, SSD is applied in three different scales for multi-modal feature extraction in our network. To strengthen local representation, we pay more attention on foreground edge and structural information by feature scaling function of SSD. Secondly, for shared feature extraction of input images and multi-modal feature fusion in all scales, we propose cross-modality feature fusion model based on SSD, consisting of Cross-Modality feature Interaction (CMI) module and Multi-Scale feature Fusion (MSF) module. CMI module is designed for cross-modality feature extraction of each scale by SSD in cross form. MSF module is designed to employ a progressive upward fusion in feature-level to obtain fine features, consisting of multi-modal features in all scales. Following coarse-to-fine, the features in 1/8 scale from CMI and 1/2 scale from MSF are collected to calculate matching probability scores. Then we respectively establish matching process by correspondences of pixel-wise. Extensive experiments demonstrate that comparing with state-of-the-art deep-learning based algorithms, RegNetMamba-2 has achieved good effects in both performance and efficiency for multi-modal image registration on the following datasets: VIS-SAR (OSDataset), VIS-IR (LGHD/RoadSence) and VIS-NIR (RGB-NIR sense).

图像配准多模态SSD特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。