融合红外与可见光图像,提升细节保留与全局特征表达
MATCNN: Infrared and Visible Image Fusion Method Based on Multi-scale CNN with Attention Transformer
- 多尺度卷积+注意力机制提取局部与全局特征
- 信息掩码增强红外目标与可见光背景纹理保留
- 在多个数据集上表现优于现有方法,适合跨模态图像融合研究
尽管基于注意力的方法在提升图像融合性能、缓解长程依赖问题方面取得显著进展,但在捕捉局部特征方面仍受限于缺乏多样化的感受野提取技术。为克服现有融合方法在多尺度局部特征提取和全局特征保持方面的不足,本文提出一种基于多尺度卷积神经网络与注意力变换器(MATCNN)的新型跨模态图像融合方法。MATCNN采用多尺度融合模块(MSFM)在不同尺度提取局部特征,利用全局特征提取模块(GFEM)捕获全局特征,二者结合可减少细节丢失并增强全局表征能力。同时引入信息掩码对图像中关键细节进行标记,旨在提升融合图像中红外显著目标及可见光背景纹理的保留比例。进一步设计了一种新型优化算法,通过整合内容损失、结构相似性指数测量与全局特征损失,利用掩码引导特征提取。在多个数据集上的定量与定性评估表明,MATCNN能有效突出红外显著目标,保留更多可见图像细节,实现更优的跨模态图像融合效果。
原文摘要 · Abstract (English)
While attention-based approaches have shown considerable progress in enhancing image fusion and addressing the challenges posed by long-range feature dependencies, their efficacy in capturing local features is compromised by the lack of diverse receptive field extraction techniques. To overcome the shortcomings of existing fusion methods in extracting multi-scale local features and preserving global features, this paper proposes a novel cross-modal image fusion approach based on a multi-scale convolutional neural network with attention Transformer (MATCNN). MATCNN utilizes the multi-scale fusion module (MSFM) to extract local features at different scales and employs the global feature extraction module (GFEM) to extract global features. Combining the two reduces the loss of detail features and improves the ability of global feature representation. Simultaneously, an information mask is used to label pertinent details within the images, aiming to enhance the proportion of preserving significant information in infrared images and background textures in visible images in fused images. Subsequently, a novel optimization algorithm is developed, leveraging the mask to guide feature extraction through the integration of content, structural similarity index measurement, and global feature loss. Quantitative and qualitative evaluations are conducted across various datasets, revealing that MATCNN effectively highlights infrared salient targets, preserves additional details in visible images, and achieves better fusion results for cross-modal images. The code of MATCNN will be available at https://github.com/zhang3849/MATCNN.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。