arXiv:2410.23905cs.CV2024-10NeurIPS被引 59

用文本控制扩散模型,实现去噪融合并突出重点物体。

Text-DiFuse: An Interactive Multi-Modal Image Fusion Framework based on Text-modulated Diffusion Model

  • 将多模态信息融合嵌入扩散过程,自适应修复图像多重退化。
  • 通过文本+零样本定位模型控制融合,显著提升目标物体清晰度。
  • 支持用户文本指令定制,适合需要精准语义融合的场景。

现有多模态图像融合方法难以应对源图像中的复合退化问题,导致融合图像存在噪声、色彩偏移、曝光不当等缺陷。同时,这些方法常忽略前景对象的特异性,削弱其在融合结果中的显著性。为此,本文提出基于文本调制扩散模型的交互式多模态图像融合框架 Text-DiFuse。首先,将特征级信息融合引入扩散过程,实现退化自适应去除与多模态信息融合,首次深度显式地将信息融合嵌入扩散机制,有效缓解复合退化问题。其次,通过将文本与零样本定位模型结合嵌入扩散融合流程,设计文本可控的融合再调节策略,实现用户自定义文本控制,增强融合图像中感兴趣对象的显著性。在多个公开数据集上的大量实验表明,Text-DiFuse 在复杂退化场景下均达到当前最优性能。此外,语义分割实验验证了该文本控制融合再调节策略对语义性能的显著提升。代码已开源:https://github.com/Leiii-Cao/Text-DiFuse。

原文摘要 · Abstract (English)

Existing multi-modal image fusion methods fail to address the compound degradations presented in source images, resulting in fusion images plagued by noise, color bias, improper exposure, \textit{etc}. Additionally, these methods often overlook the specificity of foreground objects, weakening the salience of the objects of interest within the fused images. To address these challenges, this study proposes a novel interactive multi-modal image fusion framework based on the text-modulated diffusion model, called Text-DiFuse. First, this framework integrates feature-level information integration into the diffusion process, allowing adaptive degradation removal and multi-modal information fusion. This is the first attempt to deeply and explicitly embed information fusion within the diffusion process, effectively addressing compound degradation in image fusion. Second, by embedding the combination of the text and zero-shot location model into the diffusion fusion process, a text-controlled fusion re-modulation strategy is developed. This enables user-customized text control to improve fusion performance and highlight foreground objects in the fused images. Extensive experiments on diverse public datasets show that our Text-DiFuse achieves state-of-the-art fusion performance across various scenarios with complex degradation. Moreover, the semantic segmentation experiment validates the significant enhancement in semantic performance achieved by our text-controlled fusion re-modulation strategy. The code is publicly available at https://github.com/Leiii-Cao/Text-DiFuse.

图像融合扩散模型文本控制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。