arXiv:2410.22837cs.CV2024-10中稿 · ECAI 2024被引 26

提出一种高效融合红外与可见光图像的新方法,兼顾细节与实时性。

SFDFusion: An Efficient Spatial-Frequency Domain Fusion Network for Infrared and Visible Image Fusion

  • 结合空间域与频域信息,通过双模态增强和傅里叶变换融合特征
  • 在多个公开数据集上指标领先,视觉效果清晰且目标突出
  • 结构简洁高效,适合实时视觉任务,如目标检测

红外与可见光图像融合旨在利用两种模态的互补信息,生成具有显著目标和丰富纹理细节的融合图像。现有大多数算法仅在空间域进行像素级或特征级融合,忽视了频域信息,且部分方法因结构过于复杂导致效率低下。为此,本文提出一种高效的时空频域融合网络(SFDFusion)。首先设计双模态精炼模块(DMRM),在空间域提取红外与可见光模态的互补信息并增强细粒度空间细节;其次构建频域融合模块(FDFM),通过快速傅里叶变换(FFT)将空间域转至频域,并融合频域特征;同时设计频域融合损失函数,为融合过程提供指导。大量实验表明,所提方法在多个公开数据集上各项融合指标均优于现有方法,视觉效果更优。此外,该方法兼具高效率与下游检测任务的良好表现,满足先进视觉任务的实时需求。

原文摘要 · Abstract (English)

Infrared and visible image fusion aims to utilize the complementary information from two modalities to generate fused images with prominent targets and rich texture details. Most existing algorithms only perform pixel-level or feature-level fusion from different modalities in the spatial domain. They usually overlook the information in the frequency domain, and some of them suffer from inefficiency due to excessively complex structures. To tackle these challenges, this paper proposes an efficient Spatial-Frequency Domain Fusion (SFDFusion) network for infrared and visible image fusion. First, we propose a Dual-Modality Refinement Module (DMRM) to extract complementary information. This module extracts useful information from both the infrared and visible modalities in the spatial domain and enhances fine-grained spatial details. Next, to introduce frequency domain information, we construct a Frequency Domain Fusion Module (FDFM) that transforms the spatial domain to the frequency domain through Fast Fourier Transform (FFT) and then integrates frequency domain information. Additionally, we design a frequency domain fusion loss to provide guidance for the fusion process. Extensive experiments on public datasets demonstrate that our method produces fused images with significant advantages in various fusion metrics and visual effects. Furthermore, our method demonstrates high efficiency in image fusion and good performance on downstream detection tasks, thereby satisfying the real-time demands of advanced visual tasks.

图像融合频域融合红外可见光实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。