通过频域交互提升多模态图像融合效果,更精准保留细节信息。
WIFE-Fusion:Wavelet-aware Intra-inter Frequency Enhancement for Multi-model Image Fusion
- 引入频域内自注意力与跨频交互机制,挖掘多模态特征关联。
- 在五个数据集上超越现有方法,在纹理和结构保留上表现优异。
- 适合需要高精度图像融合的医学、遥感等场景使用。
多模态图像融合能有效整合不同模态的信息,融合图像在视觉系统中至关重要。然而,现有方法常忽视频域特征探索及模态间交互关系。本文提出基于频域组件交互的多模态图像融合框架WIFE-Fusion。其核心创新包括:频域内自注意力(IFSA),通过交互式自注意力机制利用模态间的内在相关性与互补性,提取丰富的频域特征;跨频交互(IFI),通过异构频域组件间的组合交互,增强已丰富特征并过滤潜在噪声。该过程实现源特征的精确提取与特征提取-聚合的统一建模。在三个多模态融合任务的五个数据集上进行的大量实验表明,WIFE-Fusion优于当前专用与统一融合方法。代码已开源:https://github.com/Lmmh058/WIFE-Fusion。
原文摘要 · Abstract (English)
Multimodal image fusion effectively aggregates information from diverse modalities, with fused images playing a crucial role in vision systems. However, existing methods often neglect frequency-domain feature exploration and interactive relationships. In this paper, we propose wavelet-aware Intra-inter Frequency Enhancement Fusion (WIFE-Fusion), a multimodal image fusion framework based on frequency-domain components interactions. Its core innovations include: Intra-Frequency Self-Attention (IFSA) that leverages inherent cross-modal correlations and complementarity through interactive self-attention mechanisms to extract enriched frequency-domain features, and Inter-Frequency Interaction (IFI) that enhances enriched features and filters latent features via combinatorial interactions between heterogeneous frequency-domain components across modalities. These processes achieve precise source feature extraction and unified modeling of feature extraction-aggregation. Extensive experiments on five datasets across three multimodal fusion tasks demonstrate WIFE-Fusion's superiority over current specialized and unified fusion methods. Our code is available at https://github.com/Lmmh058/WIFE-Fusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。