用文本语义指导红外与可见光图像融合,提升检测分割效果
TeSG: Textual Semantic Guidance for Infrared and Visible Image Fusion
- 从大模型提取文本描述,生成掩码与语义双层引导信号
- 在下游检测与分割任务中,性能优于当前最先进方法
- 适合需要高精度多模态图像融合的应用场景
红外与可见光图像融合(IVF)旨在结合两种模态的互补信息,生成更丰富全面的输出。近年来,文本引导的IVF因其灵活性和通用性展现出巨大潜力,但文本语义信息的有效整合与利用仍不充分。为此,本文在掩码语义与文本语义两个层面引入文本语义信息,均来自大视觉-语言模型(VLMs)提取的文本描述。在此基础上,提出面向红外与可见光图像融合的文本语义引导方法TeSG,该方法可优化下游任务如目标检测与分割的表现。TeSG包含三个核心组件:语义信息生成器(SIG),用于基于文本描述生成掩码与文本语义;掩码引导交叉注意力(MGCA)模块,利用掩码语义对红外与可见光图像特征进行初始注意力融合;以及文本驱动注意力融合(TDAF)模块,通过文本语义驱动的门控注意力进一步优化融合过程。大量实验表明,本方法在下游任务中表现优异,超越现有最先进方法。
原文摘要 · Abstract (English)
Infrared and visible image fusion (IVF) aims to combine complementary information from both image modalities, producing more informative and comprehensive outputs. Recently, text-guided IVF has shown great potential due to its flexibility and versatility. However, the effective integration and utilization of textual semantic information remains insufficiently studied. To tackle these challenges, we introduce textual semantics at two levels: the mask semantic level and the text semantic level, both derived from textual descriptions extracted by large Vision-Language Models (VLMs). Building on this, we propose Textual Semantic Guidance for infrared and visible image fusion, termed TeSG, which guides the image synthesis process in a way that is optimized for downstream tasks such as detection and segmentation. Specifically, TeSG consists of three core components: a Semantic Information Generator (SIG), a Mask-Guided Cross-Attention (MGCA) module, and a Text-Driven Attentional Fusion (TDAF) module. The SIG generates mask and text semantics based on textual descriptions. The MGCA module performs initial attention-based fusion of visual features from both infrared and visible images, guided by mask semantics. Finally, the TDAF module refines the fusion process with gated attention driven by text semantics. Extensive experiments demonstrate the competitiveness of our approach, particularly in terms of performance on downstream tasks, compared to existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。