arXiv:2509.11476cs.CVcs.LG2025-09ICCV被引 8

通过目标感知监督提升红外可见光图像融合的语义保真度

Modality-Aware Infrared and Visible Image Fusion with Target-Aware Supervision

  • 引入模态感知注意力,动态调整红外与可见光特征贡献
  • 采用像素级自适应融合权重,实现细粒度可解释融合
  • 利用弱标注感兴趣区域,增强关键目标的语义一致性

红外与可见光图像融合(IVIF)是多模态感知的基础任务,旨在整合不同波段的结构与纹理信息。本文提出FusionNet,一种端到端融合框架,显式建模跨模态交互并增强任务关键区域。该框架引入模态感知注意力机制,根据特征判别能力动态调节红外与可见光特征的贡献。为实现细粒度、可解释的融合,进一步设计像素级α混合模块,以自适应且内容感知的方式学习空间变化的融合权重。此外,提出目标感知损失,利用弱标注的兴趣区域(ROI)监督,确保重要物体(如行人、车辆)所在区域的语义一致性。在公开数据集M3FD上的实验表明,FusionNet生成的融合图像具有更强的语义保留、更高的感知质量与清晰的可解释性。本框架为语义感知的多模态图像融合提供了通用且可扩展的解决方案,有助于下游任务如目标检测与场景理解。

原文摘要 · Abstract (English)

Infrared and visible image fusion (IVIF) is a fundamental task in multi-modal perception that aims to integrate complementary structural and textural cues from different spectral domains. In this paper, we propose FusionNet, a novel end-to-end fusion framework that explicitly models inter-modality interaction and enhances task-critical regions. FusionNet introduces a modality-aware attention mechanism that dynamically adjusts the contribution of infrared and visible features based on their discriminative capacity. To achieve fine-grained, interpretable fusion, we further incorporate a pixel-wise alpha blending module, which learns spatially-varying fusion weights in an adaptive and content-aware manner. Moreover, we formulate a target-aware loss that leverages weak ROI supervision to preserve semantic consistency in regions containing important objects (e.g., pedestrians, vehicles). Experiments on the public M3FD dataset demonstrate that FusionNet generates fused images with enhanced semantic preservation, high perceptual quality, and clear interpretability. Our framework provides a general and extensible solution for semantic-aware multi-modal image fusion, with benefits for downstream tasks such as object detection and scene understanding.

图像融合多模态目标感知注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。