提出频域感知多视角融合网络,提升红外可见光图像融合效果
FMRFusion: Frequency-Aware Multi-View Representation Learning for Heterogeneous Image Fusion

- 分频处理+多尺度结构感知,捕捉细节与上下文信息
- 跨视图互补交互,有效融合光照与辐射特性差异
- 在夜间等复杂场景下表现优异,适合实际应用
红外与可见光图像融合旨在生成同时保留显著目标信息和丰富纹理细节的复合图像,融合两种异质模态。以往方法多采用单模块堆叠提取特征,易导致模态特异性学习不完整,影响融合效果与真实场景下的鲁棒性。为此,本文提出FMRFusion:一种面向异质图像融合的频域感知多视角表征学习网络。引入多尺度结构感知模块,有效提取细粒度局部结构与关键上下文信息;采用双线性频域分解机制,将特征分离为高低频成分,实现不同频域下局部细节与全局表征的联合建模;设计跨视图互补交互模块,显式建模反射光信息与辐射强度响应间的互补特性,促进跨视角有效交互;进一步通过流匹配优化融合结果,逐步学习从粗略数据到高质量表示的映射。在多个基准数据集上的大量实验表明,FMRFusion在各类融合任务中均取得卓越且一致的性能,尤其在夜间场景下表现突出。
原文摘要 · Abstract (English)
Infrared and visible image fusion aims to generate a composite image that retains significant target information and preserves detailed textures, integrating two heterogeneous modalities. Previous image fusion methods typically adopt a single-module stacking approach to extract features from the two modalities. However, these approaches may result in incomplete learning of their distinct characteristics, thereby limiting the fusion effectiveness and constrain ing robustness in real-world heterogeneous data scenarios. To address these challenges, we propose FMRFusion, a frequency-aware multi-view representation learning network for Heterogeneous Image Fusion. A Multi-Scale Struc tural Perception Module is introduced to effectively capture discriminative structures, extracting fine-grained local structures and essential contextual information. A bilinear frequency decomposition mechanism is employed to sepa rate features into high-frequency and low-frequency components, enabling joint modeling of local details and global representations across different frequency domains. Moreover, a Cross-View Complementary Interaction is incorpo rated to explicitly model and fuse the complementary characteristics between reflected light information and radiative intensity responses, facilitating effective cross-view interaction. We further improve the Performance of the fused results by flow matching, which progressively refines the fused features by learning the transformation from coarse data to high-quality representations. Extensive experiments conducted on multiple benchmark datasets demonstrate that FMRFusion achieves superior and consistent performance across a range of fusion tasks, especially in nighttime scenarios
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。