arXiv:2509.05000cs.CV2025-09

用视觉语言模型感知降级,实现红外可见光图像融合的鲁棒性提升

Dual-Domain Perspective on Degradation-Aware Fusion: A VLM-Guided Robust Infrared and Visible Image Fusion Framework

  • 引入视觉语言模型感知双源降级,联合优化频域与空间域特征
  • 在降级场景下优于现有方法,显著减少误差累积与性能下降
  • 适合复杂环境下的多模态图像融合任务,如夜间监控、安防系统

现有红外-可见光图像融合(IVIF)方法通常假设输入质量良好,因此在双源降级场景下表现不佳,需手动选择并顺序应用多个预增强步骤。这种解耦的预增强-融合流程必然导致误差累积和性能下降。为此,我们提出引导式双域融合框架GD^2Fusion,将视觉语言模型(VLMs)用于降级感知,并结合频域与空间域的联合优化。具体而言,设计的引导式频率模态特异性提取(GFMSE)模块在频域实现降级感知与抑制,并判别性地提取融合相关子带特征;同时,引导式空间模态聚合融合(GSMAF)模块在空间域执行跨模态降级滤波与自适应多源特征聚合,增强模态互补性与结构一致性。大量定性与定量实验表明,相较于现有算法与策略,GD^2Fusion在双源降级场景下实现了更优的融合性能。代码将在论文被接收后公开。

原文摘要 · Abstract (English)

Most existing infrared-visible image fusion (IVIF) methods assume high-quality inputs, and therefore struggle to handle dual-source degraded scenarios, typically requiring manual selection and sequential application of multiple pre-enhancement steps. This decoupled pre-enhancement-to-fusion pipeline inevitably leads to error accumulation and performance degradation. To overcome these limitations, we propose Guided Dual-Domain Fusion (GD^2Fusion), a novel framework that synergistically integrates vision-language models (VLMs) for degradation perception with dual-domain (frequency/spatial) joint optimization. Concretely, the designed Guided Frequency Modality-Specific Extraction (GFMSE) module performs frequency-domain degradation perception and suppression and discriminatively extracts fusion-relevant sub-band features. Meanwhile, the Guided Spatial Modality-Aggregated Fusion (GSMAF) module carries out cross-modal degradation filtering and adaptive multi-source feature aggregation in the spatial domain to enhance modality complementarity and structural consistency. Extensive qualitative and quantitative experiments demonstrate that GD^2Fusion achieves superior fusion performance compared with existing algorithms and strategies in dual-source degraded scenarios. The code will be publicly released after acceptance of this paper.

图像融合视觉语言模型降级感知多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。