arXiv:2510.26268cs.CV2025-10NeurIPS被引 3

基于人脑认知规律,提升红外与可见光图像融合的结构一致性和细节质量。

Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws

  • 引入多尺度掩码调节变分瓶颈编码器,精准提取低层模态信息。
  • 融合扩散模型与物理规律,实现时变物理引导生成,提升结构感知能力。
  • 在多个数据集上表现领先,适合复杂场景下的高质量图像融合任务。

现有红外与可见光图像融合方法常面临模态信息平衡难题。生成式融合方法虽通过学习数据分布重建融合图像,但生成能力有限,且模态选择缺乏可解释性,影响复杂场景下结果的可靠性与一致性。本文受人类认知规律启发,提出新型红外-可见光图像融合方法HCLFuse。首先,研究无监督融合网络中信息映射的量化理论,设计多尺度掩码调节的变分瓶颈编码器,通过后验概率建模与信息分解,精确提取低层模态特征,支撑高保真结构细节生成。其次,将扩散模型的概率生成能力与物理规律结合,构建时变物理引导机制,自适应调节不同生成阶段,增强模型对数据内在结构的感知,降低对数据质量的依赖。实验表明,该方法在多个数据集上均达到先进水平,定性与定量评估均表现优异,显著提升语义分割指标,充分验证了其在提升结构一致性与细节质量方面的优势。

原文摘要 · Abstract (English)

Existing infrared and visible image fusion methods often face the dilemma of balancing modal information. Generative fusion methods reconstruct fused images by learning from data distributions, but their generative capabilities remain limited. Moreover, the lack of interpretability in modal information selection further affects the reliability and consistency of fusion results in complex scenarios. This manuscript revisits the essence of generative image fusion under the inspiration of human cognitive laws and proposes a novel infrared and visible image fusion method, termed HCLFuse. First, HCLFuse investigates the quantification theory of information mapping in unsupervised fusion networks, which leads to the design of a multi-scale mask-regulated variational bottleneck encoder. This encoder applies posterior probability modeling and information decomposition to extract accurate and concise low-level modal information, thereby supporting the generation of high-fidelity structural details. Furthermore, the probabilistic generative capability of the diffusion model is integrated with physical laws, forming a time-varying physical guidance mechanism that adaptively regulates the generation process at different stages, thereby enhancing the ability of the model to perceive the intrinsic structure of data and reducing dependence on data quality. Experimental results show that the proposed method achieves state-of-the-art fusion performance in qualitative and quantitative evaluations across multiple datasets and significantly improves semantic segmentation metrics. This fully demonstrates the advantages of this generative image fusion method, drawing inspiration from human cognition, in enhancing structural consistency and detail quality.

图像融合生成模型认知规律多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。