融合可见光与红外图像,提升极暗环境下的图像细节还原能力。
VIIS: Visible and Infrared Information Synthesis for Severe Low-light Image Enhancement
- 提出可见-红外信息合成新任务,同时实现模态内增强与跨模态融合。
- 在无真实标签情况下,通过图像增广设计预训练任务提升模型性能。
- 采用稀疏注意力机制增强双模态特征交互,适合夜视、安防等场景应用。
在极端低光照条件下拍摄的图像常因信息缺失而难以辨识。现有单一模态增强方法难以恢复缺乏有效信息的区域。利用对光不敏感的红外图像,可见光与红外图像融合有望揭示黑暗中隐藏的信息。然而,现有方法多关注模态间互补性,忽略模态内增强,限制了输出图像的感知质量。为此,本文提出可见与红外信息合成(VIIS)新任务,旨在同时实现信息增强与模态融合。由于该任务缺乏真实标签,我们设计基于图像增广的信息合成预训练任务(ISPT)。采用扩散模型框架,并提出基于稀疏注意力的双模态残差(SADMR)条件机制,在去噪过程中自适应地迭代融合双模态先验信息。大量实验表明,所提模型在定性和定量上均优于相关领域最先进方法,以及新设计的兼具增强与融合能力的基线模型。
原文摘要 · Abstract (English)
Images captured in severe low-light circumstances often suffer from significant information absence. Existing singular modality image enhancement methods struggle to restore image regions lacking valid information. By leveraging light-impervious infrared images, visible and infrared image fusion methods have the potential to reveal information hidden in darkness. However, they primarily emphasize inter-modal complementation but neglect intra-modal enhancement, limiting the perceptual quality of output images. To address these limitations, we propose a novel task, dubbed visible and infrared information synthesis (VIIS), which aims to achieve both information enhancement and fusion of the two modalities. Given the difficulty in obtaining ground truth in the VIIS task, we design an information synthesis pretext task (ISPT) based on image augmentation. We employ a diffusion model as the framework and design a sparse attention-based dual-modalities residual (SADMR) conditioning mechanism to enhance information interaction between the two modalities. This mechanism enables features with prior knowledge from both modalities to adaptively and iteratively attend to each modality's information during the denoising process. Our extensive experiments demonstrate that our model qualitatively and quantitatively outperforms not only the state-of-the-art methods in relevant fields but also the newly designed baselines capable of both information enhancement and fusion. The code is available at https://github.com/Chenz418/VIIS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。