通过频域建模与残差先验,提升多模态图像融合质量。
Residual Prior-driven Frequency-aware Network for Image Fusion
- 双分支结构:残差先验模块提取差异信息,频域融合模块高效建模全局特征。
- 在红外-可见光融合任务中,显著提升纹理细节和显著目标保留能力。
- 适合需要高精度图像融合的医学、遥感等视觉任务应用。
图像融合旨在整合多模态间的互补信息以生成高质量融合图像,从而提升高层视觉任务性能。尽管全局空间建模机制已展现出良好效果,但在空间域构建长距离特征依赖关系会带来巨大计算开销。此外,缺乏真实标签进一步增加了有效捕捉互补特征的难度。为此,我们提出残差先验驱动的频域感知网络(RPFNet)。该模型采用双分支特征提取框架:残差先验模块(RPM)从残差图中提取模态特异性差异信息,为融合提供互补先验;频域融合模块(FDFM)通过频域卷积实现高效全局特征建模与融合。同时,交叉促进模块(CPM)通过双向特征交互增强局部细节与全局结构的协同感知。训练过程中,引入辅助解码器与显著性结构损失,强化模型对模态差异的敏感性;结合自适应权重的频域对比损失与SSIM损失,有效约束解空间,促进局部细节与全局特征的联合捕捉,并确保互补信息保留。大量实验验证了RPFNet的融合性能,其能有效整合判别性特征,增强纹理细节与显著物体表达,可有效支持高层视觉任务部署。
原文摘要 · Abstract (English)
Image fusion aims to integrate complementary information across modalities to generate high-quality fused images, thereby enhancing the performance of high-level vision tasks. While global spatial modeling mechanisms show promising results, constructing long-range feature dependencies in the spatial domain incurs substantial computational costs. Additionally, the absence of ground-truth exacerbates the difficulty of capturing complementary features effectively. To tackle these challenges, we propose a Residual Prior-driven Frequency-aware Network, termed as RPFNet. Specifically, RPFNet employs a dual-branch feature extraction framework: the Residual Prior Module (RPM) extracts modality-specific difference information from residual maps, thereby providing complementary priors for fusion; the Frequency Domain Fusion Module (FDFM) achieves efficient global feature modeling and integration through frequency-domain convolution. Additionally, the Cross Promotion Module (CPM) enhances the synergistic perception of local details and global structures through bidirectional feature interaction. During training, we incorporate an auxiliary decoder and saliency structure loss to strengthen the model's sensitivity to modality-specific differences. Furthermore, a combination of adaptive weight-based frequency contrastive loss and SSIM loss effectively constrains the solution space, facilitating the joint capture of local details and global features while ensuring the retention of complementary information. Extensive experiments validate the fusion performance of RPFNet, which effectively integrates discriminative features, enhances texture details and salient objects, and can effectively facilitate the deployment of the high-level vision task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。