轻量级网络提升多源图像融合效率与质量
Combined Dictionary Unfolding Network with Gradient-Adaptive Fidelity for Transferable Multi-Source Fusion

- 将耦合字典学习的共性-特性分解先验转为联合展开架构
- 在TNO和RoadScene数据集上多项指标优于现有方法
- 适合边缘设备部署,无需真实图像监督训练
基于深度展开网络的方法通过结合模型驱动的迭代优化与数据驱动的深度学习,已成为多源图像融合的有效方案。然而,现有方法大多源于交替最小化,分步更新不同模态特征,导致计算与内存开销大,难以在资源受限的边缘设备上部署。为此,我们提出CDNet,一种轻量级联合字典展开网络,用于多源图像融合。不同于引入新的稀疏编码先验或经验压缩现有网络,CDNet将耦合字典学习的共性-特性分解先验转化为结构约束的联合展开架构。由此产生的CDBlock采用块稀疏交互拓扑,执行由模型导出的共性和模态特异性表示的联合更新,从而简化特征学习并提升效率。此外,我们设计了一种紧凑的高低频图像保真度损失,实现无真实图像监督下的训练。我们在四个任务上评估了CDNet:多曝光图像融合、红外与可见光图像融合、医学图像融合以及用于语义分割的红外与可见光图像融合。实验结果表明,CDNet在保持高效率的同时实现了具有竞争力或更优的融合性能。在红外与可见光图像融合任务中,于TNO数据集上六项指标中有四项优于竞品,在RoadScene数据集上五项指标优于竞品;尤其在PSNR上分别领先第二名1.23 dB和1.59 dB。
原文摘要 · Abstract (English)
Deep Unfolding Network-based methods have emerged as effective solutions for multi-source image fusion by combining model-driven iterative optimization with data-driven deep learning. However, most existing deep unfolding image fusion methods are derived from alternating minimization, which updates the features of different modalities separately. This design introduces considerable computational and memory overhead, limiting deployment on resource-constrained edge devices. To address this issue, we propose CDNet, a lightweight Combined Dictionary Unfolding Network for multi-source image fusion. Rather than introducing a new sparse coding prior or empirically compressing an existing fusion network, CDNet translates the unique-common decomposition prior of coupled dictionary learning into a structurally constrained joint unfolding architecture. The resulting CDBlock follows a block-sparse interaction topology and performs a model-derived joint update of common and modality-specific representations, thereby streamlining feature learning and improving efficiency.In addition, we design a compact High- and Low-frequency Image Fidelity loss for unsupervised training without ground-truth images. We evaluate CDNet on four tasks, including multi-exposure image fusion, infrared and visible image fusion, medical image fusion, and infrared and visible image fusion for semantic segmentation. Experimental results show that CDNet achieves competitive or superior fusion performance with high efficiency. For infrared and visible image fusion, CDNet outperforms competing methods on four of six metrics on the TNO dataset and five of six metrics on the RoadScene dataset. In particular, it surpasses the second-best method by 1.23 dB and 1.59 dB in PSNR on TNO and RoadScene, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。