arXiv:2503.18627cs.CVcs.AI2025-03被引 1

提出动态融合机制,让扩散模型根据阶段变化自动调整多源图像贡献。

Dig2DIG: Dig into Diffusion Information Gains for Image Fusion

  • 基于去噪过程中的动态信息增益,实现模态权重自适应
  • 理论证明可降低泛化误差上界,提升融合可靠性
  • 在多种场景下兼顾质量与速度,适合实时图像融合应用

图像融合通过整合多源图像的互补信息生成更丰富结果。近期扩散模型因强大的生成能力被引入图像融合,但现有方法通常采用预设的多模态引导,无法捕捉模态重要性的动态变化,且缺乏理论保障。本文揭示了去噪过程中存在显著的时空不平衡性:扩散模型在不同图像区域和去噪步骤中产生动态信息增益。基于此,我们提出Dig2DIG框架,理论上推导出一种基于扩散的信息增益动态融合机制,可证明地降低泛化误差上界。进一步引入扩散信息增益(DIG)量化各模态在不同去噪阶段的贡献,实现动态引导。大量实验表明,该方法在多种融合场景下均优于现有扩散基方法,在融合质量与推理效率方面均有提升。

原文摘要 · Abstract (English)

Image fusion integrates complementary information from multi-source images to generate more informative results. Recently, the diffusion model, which demonstrates unprecedented generative potential, has been explored in image fusion. However, these approaches typically incorporate predefined multimodal guidance into diffusion, failing to capture the dynamically changing significance of each modality, while lacking theoretical guarantees. To address this issue, we reveal a significant spatio-temporal imbalance in image denoising; specifically, the diffusion model produces dynamic information gains in different image regions with denoising steps. Based on this observation, we Dig into the Diffusion Information Gains (Dig2DIG) and theoretically derive a diffusion-based dynamic image fusion framework that provably reduces the upper bound of the generalization error. Accordingly, we introduce diffusion information gains (DIG) to quantify the information contribution of each modality at different denoising steps, thereby providing dynamic guidance during the fusion process. Extensive experiments on multiple fusion scenarios confirm that our method outperforms existing diffusion-based approaches in terms of both fusion quality and inference efficiency.

图像融合扩散模型动态融合信息增益

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。