通过傅里叶变换与多尺度注意力融合,提升低光图像细节恢复能力。
Low-light Image Enhancement via Multi-scale Attention combined with Fourier Transform

- 结合傅里叶域与多尺度注意力,从频域增强图像特征。
- 在SDSD户外数据集上,PSNR达41.76 dB,比Retinexformer提升11.92 dB。
- 适合需要高保真纹理恢复的低光图像处理场景。
低光图像增强(LLIE)旨在改善复杂光照条件下的图像质量与清晰度。现有基于深度学习的方法难以准确捕捉真实光照并还原纹理细节,主要因算法优势未被充分挖掘。为此,本文提出一种监督式频域深度网络MSFT,采用一阶段U形架构,通过多尺度注意力将低光图像信息引入网络。进一步在自创模块中融合先验通道的幅值信息与低光图像幅值,并实现多尺度引导。为更好增强微弱特征(如细粒度内容与纹理),并在解码阶段融合全局上下文置信度,我们分别引入多形状协同注意力与轻量级网络,有效整合高维空间信息,嵌入富含纹理的超优特征空间。在LOL、SID、SMID和SDSD数据集上的大量实验表明,MSFT显著优于当前最优方法。例如,在SDSD户外数据集上,相比Retinexformer,本方法峰值信噪比(PSNR)达到41.76分贝,提升11.92分贝,结构相似性指数(SSIM)达0.988,提高13.80%。
原文摘要 · Abstract (English)
Low-light image enhancement (LLIE) aims to improve image quality and clarity in diverse and demanding low-illumination environments. However, existing deep learning-based LLIE methods struggle to accurately capture real-world illumination and restore texture details, largely because their algorithmic strengths remain underutilized. To address these issues, we present a supervised frequency domain deep learning network for LLIE, named multi-scale attention combined with the Fourier transform (MSFT) which adopts a U-shaped, one-stage architecture that infuses guidance from low-light images into the network by channeling it through multi-scale attention. We further fuse the amplitude information from priori channels with that of the low-light image in MSFT's self-created module, and carry out multi-scale guidance along with the network. Subsequently, to better enhance the faint feature, such as fine content and textures, and to better fuse global context confidence in the decoding stage, we separately introduce a multi-shape synergistic attention and a lightweight network that effectively integrate information in high-dimensional space to embed into the superlative feature space channel containing rich texture information. Extensive experiments conducted on LOL, SID, SMID, and SDSD datasets demonstrate that MSFT significantly outperforms state-of-the-art competitors. For example, compared with Retinexformer, our method achieves a peak signal-to-noise ratio of up to 41.76 decibels on the SDSD-outdoor dataset with an increase of 11.92 decibels and a structural similarity index of 0.988 with a 13.80% improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。