arXiv:2601.01870cs.CV2026-01中稿 · IEEE Transactions …被引 1

用实体级文本指导红外可见光图像融合,提升语义准确性

Entity-Guided Multi-Task Learning for Infrared and Visible Image Fusion

  • 从视觉语言模型生成的描述中提取实体级文本,去除冗余噪声
  • 多任务学习框架联合图像融合与多标签分类,提升融合质量
  • 实体引导跨模态交互模块,增强视觉与语义特征的细粒度对齐

现有基于文本的红外与可见光图像融合方法通常依赖句子级文本信息,易引入语义噪声且未能充分挖掘文本深层语义价值。为此,提出一种名为实体引导多任务学习(EGMT)的新融合方法。该方法包含三个创新组件:(i) 提出一种合理方法,从大视觉-语言模型生成的图像描述中提取实体级文本信息,消除原始文本中的语义噪声,同时保留关键语义;(ii) 构建并行多任务学习架构,将图像融合与多标签分类任务结合,利用实体作为伪标签提供语义监督,使模型更深入理解图像内容,显著提升融合图像的质量与语义密度;(iii) 设计实体引导的跨模态交互模块,促进视觉与实体级文本特征之间的细粒度交互,通过捕捉跨模态依赖关系,在视觉内部与视觉-实体层面增强特征表示。为推动该框架的广泛应用,发布四个公开数据集(TNO、RoadScene、M3FD、MSRS)的实体标注版本。大量实验表明,EGMT在保持显著目标、纹理细节和语义一致性方面优于当前最优方法。代码与数据集将公开于 https://github.com/wyshao-01/EGMT。

原文摘要 · Abstract (English)

Existing text-driven infrared and visible image fusion approaches often rely on textual information at the sentence level, which can lead to semantic noise from redundant text and fail to fully exploit the deeper semantic value of textual information. To address these issues, we propose a novel fusion approach named Entity-Guided Multi-Task learning for infrared and visible image fusion (EGMT). Our approach includes three key innovative components: (i) A principled method is proposed to extract entity-level textual information from image captions generated by large vision-language models, eliminating semantic noise from raw text while preserving critical semantic information; (ii) A parallel multi-task learning architecture is constructed, which integrates image fusion with a multi-label classification task. By using entities as pseudo-labels, the multi-label classification task provides semantic supervision, enabling the model to achieve a deeper understanding of image content and significantly improving the quality and semantic density of the fused image; (iii) An entity-guided cross-modal interactive module is also developed to facilitate the fine-grained interaction between visual and entity-level textual features, which enhances feature representation by capturing cross-modal dependencies at both inter-visual and visual-entity levels. To promote the wide application of the entity-guided image fusion framework, we release the entity-annotated version of four public datasets (i.e., TNO, RoadScene, M3FD, and MSRS). Extensive experiments demonstrate that EGMT achieves superior performance in preserving salient targets, texture details, and semantic consistency, compared to the state-of-the-art methods. The code and dataset will be publicly available at https://github.com/wyshao-01/EGMT.

图像融合多模态实体引导跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。