arXiv:2503.18378cs.CV2025-03被引 6

用小波变换与状态空间模型融合红外可见光图像,提升细节保留和全局特征提取。

Exploring State Space Model in Wavelet Domain: An Infrared and Visible Image Fusion Network via Wavelet Transform and State Space Model

  • 结合小波变换与状态空间模型,同时捕捉频率域特征与全局信息。
  • 在多个数据集上优于当前最佳方法,峰值信噪比最高提升0.47dB。
  • 适合需要高保真多模态图像融合的遥感、安防等场景。

深度学习技术已显著推动红外与可见光图像融合(IVIF)发展,但在复杂场景中仍存在未能充分融合频域特征与全局语义信息的问题,导致跨模态全局特征提取不充分及局部纹理细节丢失。为此,本文提出Wavelet-Mamba(W-Mamba),将小波变换与状态空间模型(SSM)相结合。设计了基于小波的频域特征提取模块与利用SSM实现全局信息捕获的Wavelet-SSM模块,有效融合局部与全局特征。进一步提出跨模态特征注意力调制机制,促进不同模态间高效交互与融合。实验表明,该方法在视觉效果与定量指标上均优于现有先进方法,在IRST-2019、FLIR-AVIRIS等数据集上平均峰值信噪比(PSNR)提升0.18–0.47dB,且计算开销更低。代码已开源。

原文摘要 · Abstract (English)

Deep learning techniques have revolutionized the infrared and visible image fusion (IVIF), showing remarkable efficacy on complex scenarios. However, current methods do not fully combine frequency domain features with global semantic information, which will result in suboptimal extraction of global features across modalities and insufficient preservation of local texture details. To address these issues, we propose Wavelet-Mamba (W-Mamba), which integrates wavelet transform with the state-space model (SSM). Specifically, we introduce Wavelet-SSM module, which incorporates wavelet-based frequency domain feature extraction and global information extraction through SSM, thereby effectively capturing both global and local features. Additionally, we propose a cross-modal feature attention modulation, which facilitates efficient interaction and fusion between different modalities. The experimental results indicate that our method achieves both visually compelling results and superior performance compared to current state-of-the-art methods. Our code is available at https://github.com/Lmmh058/W-Mamba.

图像融合小波变换状态空间模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。