arXiv:2509.12901cs.CV2025-09

用结构化场景图提升红外可见光图像融合的语义与细节表现

MSGFusion: Multimodal Scene Graph-Guided Infrared and Visible Image Fusion

  • 通过文本和视觉构建结构化场景图,显式建模实体与关系
  • 在多个数据集上优于当前最佳方法,尤其在细节保留和结构清晰度上
  • 适合需要高语义一致性的多模态图像融合任务

红外与可见光图像融合因在复杂恶劣环境中的强互补性受到广泛关注。尽管基于深度学习的方法在特征提取、对齐、融合与重建方面取得显著进展,但仍严重依赖纹理、对比度等低层视觉线索,难以捕捉图像中的高层语义信息。近期尝试引入文本作为语义引导,但依赖非结构化描述,未显式建模实体、属性与关系,也缺乏空间定位,限制了细粒度融合性能。为此,本文提出MSGFusion,一种多模态场景图引导的红外与可见光图像融合框架。通过深度耦合来自文本和视觉的结构化场景图,显式表示实体、属性与空间关系,并通过逐级模块同步优化高层语义与低层细节,包括场景图表示、分层聚合与图驱动融合。在多个公开基准上的大量实验表明,MSGFusion显著优于现有先进方法,尤其在细节保留与结构清晰度方面表现优异,并在低光目标检测、语义分割和医学图像融合等下游任务中展现出更强的语义一致性与泛化能力。

原文摘要 · Abstract (English)

Infrared and visible image fusion has garnered considerable attention owing to the strong complementarity of these two modalities in complex, harsh environments. While deep learning-based fusion methods have made remarkable advances in feature extraction, alignment, fusion, and reconstruction, they still depend largely on low-level visual cues, such as texture and contrast, and struggle to capture the high-level semantic information embedded in images. Recent attempts to incorporate text as a source of semantic guidance have relied on unstructured descriptions that neither explicitly model entities, attributes, and relationships nor provide spatial localization, thereby limiting fine-grained fusion performance. To overcome these challenges, we introduce MSGFusion, a multimodal scene graph-guided fusion framework for infrared and visible imagery. By deeply coupling structured scene graphs derived from text and vision, MSGFusion explicitly represents entities, attributes, and spatial relations, and then synchronously refines high-level semantics and low-level details through successive modules for scene graph representation, hierarchical aggregation, and graph-driven fusion. Extensive experiments on multiple public benchmarks show that MSGFusion significantly outperforms state-of-the-art approaches, particularly in detail preservation and structural clarity, and delivers superior semantic consistency and generalizability in downstream tasks such as low-light object detection, semantic segmentation, and medical image fusion.

图像融合多模态场景图红外可见光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。