用大模型理解天气退化,一次解决雾霾雨雪下的红外可见光融合。
MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics
- 通过视觉语言模型提取天气退化语义先验,指导融合过程。
- 引入专家混合系统,支持雾霾、雨、雪等多类退化场景统一处理。
- 在复杂天气下优于现有方法,适合实际户外成像应用。
红外与可见光图像融合旨在将多模态互补信息整合为单一结果。然而,现有方法未能考虑恶劣天气下可见光图像的退化,导致融合性能下降;且依赖固定网络结构,难以适应多样退化场景。为此,我们提出一种基于大语言模型的统一退化感知图像融合框架(MdaIF)。针对不同退化场景(如雾霾、雨、雪)在大气传输中的散射特性差异,引入混合专家(MoE)系统以应对多退化场景。通过预训练视觉-语言模型(VLM)提取天气感知的退化知识与场景特征表示,统称为语义先验。在此基础上,提出退化感知通道注意力模块(DCAM),利用退化原型分解促进通道域内的多模态特征交互。此外,结合语义先验与通道调制特征实现有效专家路由,提升复杂退化场景下的融合鲁棒性。大量实验验证了MdaIF的有效性,在多个基准数据集上显著优于当前最优方法。
原文摘要 · Abstract (English)
Infrared and visible image fusion aims to integrate complementary multi-modal information into a single fused result. However, existing methods 1) fail to account for the degradation visible images under adverse weather conditions, thereby compromising fusion performance; and 2) rely on fixed network architectures, limiting their adaptability to diverse degradation scenarios. To address these issues, we propose a one-stop degradation-aware image fusion framework for multi-degradation scenarios driven by a large language model (MdaIF). Given the distinct scattering characteristics of different degradation scenarios (e.g., haze, rain, and snow) in atmospheric transmission, a mixture-of-experts (MoE) system is introduced to tackle image fusion across multiple degradation scenarios. To adaptively extract diverse weather-aware degradation knowledge and scene feature representations, collectively referred to as the semantic prior, we employ a pre-trained vision-language model (VLM) in our framework. Guided by the semantic prior, we propose degradation-aware channel attention module (DCAM), which employ degradation prototype decomposition to facilitate multi-modal feature interaction in channel domain. In addition, to achieve effective expert routing, the semantic prior and channel-domain modulated features are utilized to guide the MoE, enabling robust image fusion in complex degradation scenarios. Extensive experiments validate the effectiveness of our MdaIF, demonstrating superior performance over SOTA methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。