arXiv:2509.09427cs.CV2025-09被引 32

融合多模态图像并提升分辨率,同时增强语义信息。

FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

  • 将融合与超分统一为条件生成问题,引入语义引导机制。
  • 在6个数据集上实现更优细节恢复,最高提升0.39dB PSNR。
  • 适合军事侦察、远距离探测等需高清晰度的场景。

图像融合是重要的信息融合与低层视觉技术,通过整合源图像的互补信息生成信息丰富的融合图像。近年来已有研究尝试联合实现图像融合与超分辨率,但在军事侦察、远距离探测等真实场景中,多模态图像的目标与背景结构常因分辨率低、语义信息弱而受损,导致现有方法效果不佳。为此,我们提出FS-Diff:一种语义引导与清晰度感知的联合图像融合与超分辨率方法。该方法将融合与超分统一为条件生成任务,利用提出的清晰度感知机制提供语义引导,实现自适应低分辨率感知与跨模态特征提取。具体地,将期望的融合结果初始化为纯高斯噪声,并引入双向特征Mamba网络以提取多模态图像的全局特征。进一步,以源图像和语义作为条件,通过改进的U-Net网络执行随机迭代去噪过程,该网络在多个噪声水平下进行训练,从而生成具有跨模态特征与丰富语义信息的高分辨率融合图像。我们还构建了覆盖600对图像的航空视图多场景(AVMS)基准数据集。在六个公开及自建数据集上的大量联合融合与超分辨率实验表明,FS-Diff在多种放大倍数下均优于现有最先进方法,能更好恢复细节与语义信息。代码已开源:https://github.com/XylonXu01/FS-Diff。

原文摘要 · Abstract (English)

As an influential information fusion and low-level vision technique, image fusion integrates complementary information from source images to yield an informative fused image. A few attempts have been made in recent years to jointly realize image fusion and super-resolution. However, in real-world applications such as military reconnaissance and long-range detection missions, the target and background structures in multimodal images are easily corrupted, with low resolution and weak semantic information, which leads to suboptimal results in current fusion techniques. In response, we propose FS-Diff, a semantic guidance and clarity-aware joint image fusion and super-resolution method. FS-Diff unifies image fusion and super-resolution as a conditional generation problem. It leverages semantic guidance from the proposed clarity sensing mechanism for adaptive low-resolution perception and cross-modal feature extraction. Specifically, we initialize the desired fused result as pure Gaussian noise and introduce the bidirectional feature Mamba to extract the global features of the multimodal images. Moreover, utilizing the source images and semantics as conditions, we implement a random iterative denoising process via a modified U-Net network. This network istrained for denoising at multiple noise levels to produce high-resolution fusion results with cross-modal features and abundant semantic information. We also construct a powerful aerial view multiscene (AVMS) benchmark covering 600 pairs of images. Extensive joint image fusion and super-resolution experiments on six public and our AVMS datasets demonstrated that FS-Diff outperforms the state-of-the-art methods at multiple magnifications and can recover richer details and semantics in the fused images. The code is available at https://github.com/XylonXu01/FS-Diff.

图像融合超分辨率多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。