统一扩散模型解决多任务图像融合,抗退化能力强。
UniDiffFusion: A Unified Diffusion Framework for Multi-Task and Degradation-Robust Image Fusion

- 用预训练扩散模型构建共享融合主干,支持多任务
- 动态路由退化提示,恢复受损特征后融合
- 适配检测/分割等下游任务,无需改动主干
通用图像融合旨在整合多源图像的互补信息,但现有方法多依赖特定任务模型,且在不同退化条件下性能不稳定。本文提出UniDiffFusion,一种统一的扩散框架,实现多任务与退化鲁棒的图像融合。该框架利用预训练扩散模型的强大生成先验,建立跨异构任务的共享融合主干,并引入任务与退化感知的条件自适应机制,以满足不同信息选择需求。具体地,采用任务提示调制逐步适应共享扩散表示至不同融合目标;设计退化提示路由器,动态检索退化感知先验,在融合前恢复被污染的源特征。此外,引入应用提示库,融入面向任务的语义引导(如目标检测、语义分割),而无需修改共享融合与恢复路径。框架采用渐进式训练,解耦融合学习、退化感知恢复与应用特异性适应,降低异构目标间的干扰。在可见光-红外、多曝光、多聚焦图像融合任务上的大量实验表明,UniDiffFusion在清洁与退化条件下均取得更优融合质量与鲁棒性。同时,显著提升下游检测与语义分割性能,验证其作为统一扩散框架在感知融合与任务导向视觉中的有效性。
原文摘要 · Abstract (English)
General image fusion aims to integrate complementary information from multiple source images, but existing methods often rely on task-specific models and struggle to maintain robust performance under diverse degradation conditions. In this paper, we propose UniDiffFusion, a unified diffusion framework for multi-task and degradation-robust image fusion. UniDiffFusion leverages the strong generative prior of a pretrained diffusion model to establish a shared fusion backbone across heterogeneous fusion tasks, while introducing task- and degradation-aware conditional adaptation to accommodate their distinct information-selection requirements. Specifically, we employ task prompt modulation to progressively adapt the shared diffusion representations to different fusion objectives, and develop a degradation prompt router to dynamically retrieve degradation-aware priors and restore corrupted source features before fusion. Furthermore, an application prompt bank is introduced to incorporate task-oriented semantic guidance for downstream applications, such as object detection and semantic segmentation, without altering the shared fusion and restoration pathways. The proposed framework is trained in a progressive manner to decouple fusion learning, degradation-aware restoration, and application-specific adaptation, thereby reducing interference among heterogeneous objectives. Extensive experiments on visible-infrared, multi-exposure, and multi-focus image fusion demonstrate that UniDiffFusion achieves superior fusion quality and robustness under both clean and degraded conditions. Moreover, UniDiffFusion consistently improves downstream detection and semantic segmentation performance, demonstrating its effectiveness as a unified diffusion framework for both perceptual fusion and task-oriented vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。