arXiv:2508.05264cs.CVcs.AI2025-08被引 17

用SAM引导扩散模型,实现热成像与可见光图像的高保真融合。

SGDFuse: SAM-Guided Diffusion Model for High-Fidelity Infrared and Visible Image Fusion

论文配图:SGDFuse: SAM-Guided Diffusion Model for High-Fidelity Infrared and Visible Image Fusion
图 1 · 摘自论文原文
  • 用SAM提供语义先验,指导扩散模型生成
  • 两阶段设计:先结构融合,再语义引导精修
  • 显著减少伪影,提升下游任务性能

红外与可见光图像融合(IVIF)对整合热敏感信息与纹理细节至关重要,但现有方法常因‘语义盲区’导致热目标被错误抑制并引入视觉伪影。为此,我们提出基于SAM引导的扩散融合网络(SGDFuse),将IVIF重构为语义驱动的生成任务。该方法通过耦合段落任意模型(SAM)的高层语义先验与条件扩散模型的高保真生成能力,采用两阶段策略:第一阶段通过初步融合建立稳健结构基础;第二阶段利用双模态语义掩码作为空间锚点,引导扩散过程实现语义一致、高保真的重建。大量实验表明,SGDFuse不仅在图像质量上达到当前最优,还显著提升下游任务表现,验证了其作为语义感知图像融合新范式的有效性。代码已开源于 https://github.com/boshizhang123/SGDFuse。

原文摘要 · Abstract (English)

Infrared and visible image fusion (IVIF) is essential for integrating thermal saliency with textural details to support downstream perception. However, most existing approaches suffer from "semantic blindness," leading to the erroneous suppression of thermal targets and the introduction of visual artifacts. To address this, we propose SAM-Guided Diffusion Fusion Network (SGDFuse), a novel Semantic-Guided Generation (SGG) framework that reframes IVIF as a semantically-steered generative task rather than simplistic pixel mapping. Our method uniquely couples high-level semantic priors from the Segment Anything Model (SAM) with the high-fidelity generative power of a conditional diffusion model. We employ a deliberate two-stage strategy to decouple multimodal alignment from iterative refinement: Stage I establishes a robust structural foundation via preliminary fusion, while Stage II utilizes dual-modality semantic masks as spatial anchors to guide the diffusion process toward a semantically coherent, high-fidelity reconstruction. Comprehensive experiments demonstrate that SGDFuse not only delivers state-of-the-art image quality but also enhances downstream task performance, confirming its effectiveness as a new Methodological Framework for semantically aware image fusion. The code is available at https://github.com/boshizhang123/SGDFuse.

图像融合扩散模型SAM热成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。