用流形学习提升红外可见光图像融合效果
SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion
- 基于对称正定流形构建多模态图像融合网络
- 在多个公开数据集上超越现有最先进方法
- 适合做跨模态图像融合与医学影像处理的研究者
欧几里得表示学习在图像融合任务中取得了良好效果,因其在处理线性空间方面具有优势。然而,真实场景采集的数据通常具有非欧几里得结构,使用欧几里得距离评估双视图潜在表示的一致性面临挑战。为此,本文提出一种新型的对称正定(SPD)流形学习方法,用于多模态图像融合,命名为SMLNet,将图像融合方法从欧几里得空间扩展到SPD流形。具体而言,我们依据黎曼几何编码图像,以挖掘其内在统计相关性,更符合人类视觉感知。SPD矩阵是网络学习过程的核心基础。在此基础上,采用跨模态融合策略捕捉模态特异性依赖关系并增强互补信息。为在图像内在空间中捕获语义相似性,进一步设计了一个注意力模块,精细处理跨模态语义亲和矩阵。基于此,构建了基于跨模态流形学习的端到端融合网络。在多个公开数据集上的大量实验表明,该框架性能优于当前最先进方法。代码将公开于https://github.com/Shaoyun2023。
原文摘要 · Abstract (English)
Euclidean representation learning methods have achieved promising results in image fusion tasks, which can be attributed to their clear advantages in handling with linear space. However, data collected from a realistic scene usually has a non-Euclidean structure, evaluating the consistency of latent representations from paired views using Euclidean distance raises challenges. To address this issue, a novel SPD (symmetric positive definite) manifold learning is proposed for multi-modal image fusion, named SMLNet, which extends the image fusion approach from the Euclidean space to the SPD manifolds. Specifically, we encode images according to the Riemannian geometry to exploit their intrinsic statistical correlations, thereby aligning with human visual perception. The SPD matrix fundamentally underpins our network's learning process. Building upon this mathematical foundation, we employ a cross-modal fusion strategy to exploit modality-specific dependencies and augment complementary information. To capture semantic similarity in images' intrinsic space, we further develop an attention module that meticulously processes the cross-modal semantic affinity matrix. Based on this, we design an end-to-end fusion network based on cross-modal manifold learning. Extensive experiments on public datasets demonstrate that our framework exhibits superior performance compared to the current state-of-the-art methods. Our code will be publicly available at https://github.com/Shaoyun2023.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。