arXiv:2505.04526cs.CVcs.AI2025-05被引 7

解决暗光下可见与红外图像融合模糊问题,一次完成去暗、解耦与融合。

DFVO: Learning Darkness-free Visible and Infrared Image Disentanglement and Fusion All at Once

  • 采用级联多任务架构,同步实现去暗、特征解耦与图像融合。
  • 在LLVIP数据集上达到63.258 dB PSNR和0.724 CC,显著提升暗光效果。
  • 适合自动驾驶等需强光照适应能力的视觉任务应用。

可见光与红外图像融合是图像融合领域的重要任务,旨在为高层视觉任务生成具有清晰结构信息和高质量纹理特征的融合图像。然而,当可见光图像存在严重光照退化时,现有方法的融合结果常呈现模糊、昏暗的问题,严重影响自动驾驶等应用。为此,本文提出一种黑暗无损网络(DFVO),实现可见光与红外图像的解耦与融合一体化处理。该方法采用级联多任务策略,替代传统的两阶段训练(增强+融合),避免层级传输导致的信息熵损失。具体而言,构建潜在共现特征提取器(LCFE)以支持级联任务;首先设计细节提取模块(DEM)获取高频语义信息;其次引入超交叉注意力模块(HCAM)提取低频信息并保留源图像纹理特征;最后设计相关损失函数引导整体网络学习,实现更优融合效果。大量实验表明,所提方法在定性与定量评估上均优于现有先进方法。尤其在暗光环境下,可生成更清晰、信息更丰富且光照更均匀的融合结果,在LLVIP数据集上取得63.258 dB PSNR和0.724 CC的最佳表现,为高层视觉任务提供更有效信息。

原文摘要 · Abstract (English)

Visible and infrared image fusion is one of the most crucial tasks in the field of image fusion, aiming to generate fused images with clear structural information and high-quality texture features for high-level vision tasks. However, when faced with severe illumination degradation in visible images, the fusion results of existing image fusion methods often exhibit blurry and dim visual effects, posing major challenges for autonomous driving. To this end, a Darkness-Free network is proposed to handle Visible and infrared image disentanglement and fusion all at Once (DFVO), which employs a cascaded multi-task approach to replace the traditional two-stage cascaded training (enhancement and fusion), addressing the issue of information entropy loss caused by hierarchical data transmission. Specifically, we construct a latent-common feature extractor (LCFE) to obtain latent features for the cascaded tasks strategy. Firstly, a details-extraction module (DEM) is devised to acquire high-frequency semantic information. Secondly, we design a hyper cross-attention module (HCAM) to extract low-frequency information and preserve texture features from source images. Finally, a relevant loss function is designed to guide the holistic network learning, thereby achieving better image fusion. Extensive experiments demonstrate that our proposed approach outperforms state-of-the-art alternatives in terms of qualitative and quantitative evaluations. Particularly, DFVO can generate clearer, more informative, and more evenly illuminated fusion results in the dark environments, achieving best performance on the LLVIP dataset with 63.258 dB PSNR and 0.724 CC, providing more effective information for high-level vision tasks. Our code is publicly accessible at https://github.com/DaVin-Qi530/DFVO.

图像融合暗光处理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。