用大模型引导融合,让红外可见光图像既清晰又利于后续任务
OCCO: LVM-guided Infrared and Visible Image Fusion Framework based on Object-aware and Contextual COntrastive Learning
- 借助大视觉模型提供语义指导,强化对比学习捕捉关键特征
- 在四个数据集上优于8种顶尖方法,下游任务性能显著提升
- 适合需要高质量融合图像与强下游表现的跨模态应用
图像融合是计算机视觉中的关键技术,旨在生成高质量融合图像并提升下游任务表现。然而,现有方法难以兼顾二者:融合质量高时下游性能可能下降,反之亦然。为此,提出一种基于大视觉模型(LVM)引导、面向目标与上下文对比学习的新型融合框架OCCO。预训练的LVM提供语义指导,使网络专注融合任务,并通过对比学习强调显著语义特征。同时设计新型特征交互融合网络,解决因模态差异导致的信息冲突问题。通过在潜在特征空间(上下文空间)中学习正样本与负样本的区分性,提升了融合图像中目标信息的完整性,从而增强下游任务表现。在四个数据集上与八种先进方法对比,验证了该方法的有效性,且在下游视觉任务中展现出卓越性能。
原文摘要 · Abstract (English)
Image fusion is a crucial technique in the field of computer vision, and its goal is to generate high-quality fused images and improve the performance of downstream tasks. However, existing fusion methods struggle to balance these two factors. Achieving high quality in fused images may result in lower performance in downstream visual tasks, and vice versa. To address this drawback, a novel LVM (large vision model)-guided fusion framework with Object-aware and Contextual COntrastive learning is proposed, termed as OCCO. The pre-trained LVM is utilized to provide semantic guidance, allowing the network to focus solely on fusion tasks while emphasizing learning salient semantic features in form of contrastive learning. Additionally, a novel feature interaction fusion network is also designed to resolve information conflicts in fusion images caused by modality differences. By learning the distinction between positive samples and negative samples in the latent feature space (contextual space), the integrity of target information in fused image is improved, thereby benefiting downstream performance. Finally, compared with eight state-of-the-art methods on four datasets, the effectiveness of the proposed method is validated, and exceptional performance is also demonstrated on downstream visual task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。