提出差异驱动的跨模态图像融合模型,兼顾红外与可见光特征
DIFF-MF: A Difference-Driven Channel-Spatial State Space Model for Multi-Modal Image Fusion

- 用模态差异图引导特征提取,实现跨通道与空间融合
- 在驾驶与低空无人机数据集上优于现有方法,视觉与定量指标双优
- 适合需要高保真多模态图像融合的应用场景
多模态图像融合旨在整合多源图像的互补信息,生成内容更丰富的高质量融合图像。尽管基于状态空间模型的方法已实现良好性能和高计算效率,但往往过度强调红外强度而损失可见细节,或保留可见结构却削弱热目标显著性。为此,我们提出DIFF-MF,一种新型差异驱动的通道-空间状态空间模型用于多模态图像融合。该方法利用模态间特征差异图指导特征提取,并在通道与空间维度进行融合。在通道维度,通过交叉注意力双状态空间建模的通道交换模块增强通道交互,实现自适应特征重加权;在空间维度,空间交换模块采用跨模态状态空间扫描完成全面空间融合。通过高效捕捉全局依赖并保持线性计算复杂度,DIFF-MF有效集成互补多模态特征。在驾驶场景与低空无人机数据集上的实验结果表明,本方法在视觉质量与量化评估上均优于现有方法。
原文摘要 · Abstract (English)
Multi-modal image fusion aims to integrate complementary information from multiple source images to produce high-quality fused images with enriched content. Although existing approaches based on state space model have achieved satisfied performance with high computational efficiency, they tend to either over-prioritize infrared intensity at the cost of visible details, or conversely, preserve visible structure while diminishing thermal target salience. To overcome these challenges, we propose DIFF-MF, a novel difference-driven channel-spatial state space model for multi-modal image fusion. Our approach leverages feature discrepancy maps between modalities to guide feature extraction, followed by a fusion process across both channel and spatial dimensions. In the channel dimension, a channel-exchange module enhances channel-wise interaction through cross-attention dual state space modeling, enabling adaptive feature reweighting. In the spatial dimension, a spatial-exchange module employs cross-modal state space scanning to achieve comprehensive spatial fusion. By efficiently capturing global dependencies while maintaining linear computational complexity, DIFF-MF effectively integrates complementary multi-modal features. Experimental results on the driving scenarios and low-altitude UAV datasets demonstrate that our method outperforms existing approaches in both visual quality and quantitative evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。