提出双流Transformer网络,提升暗光图像增强的细节与光照不变性。
DST-Net: A Dual-Stream Transformer with Illumination-Independent Feature Guidance and Multi-Scale Spatial Convolution for Low-Light Image Enhancement
- 采用光照无关特征引导,融合DoG与VGG提取纹理先验。
- 在LOL数据集上达到25.64dB PSNR,显著提升视觉质量。
- 适合需要保留精细结构的暗光图像处理场景。
暗光图像增强旨在通过解决亮度衰减和结构破坏等信号退化问题,恢复视觉传感器在低照度环境下捕获图像的可见性。尽管已有多种算法尝试提升图像质量,现有方法常导致内在信号先验的严重丢失。为此,我们提出基于光照无关信号先验引导和多尺度空间卷积的双流Transformer网络(DST-Net)。首先,设计特征提取模块,结合高斯差分(DoG)、LAB颜色空间变换与VGG-16进行纹理提取,利用解耦的光照无关特征作为信号先验,持续指导增强过程。其次,构建双流交互架构,通过跨模态注意力机制,动态校正增强图像的退化信号表示,实现基于可微曲线估计的迭代增强。此外,为克服现有方法难以保留细粒度结构与纹理的问题,提出多尺度空间融合块(MSFB),融合伪3D与3D梯度算子卷积,显式恢复高频边缘,同时通过多尺度空间卷积捕捉通道间空间相关性。大量实验与消融研究证明,DST-Net在主观视觉质量和客观指标上均表现优异。具体而言,在LOL数据集上取得25.64 dB PSNR;在LSRW数据集上的验证进一步证实其跨场景鲁棒泛化能力。
原文摘要 · Abstract (English)
Low-light image enhancement aims to restore the visibility of images captured by visual sensors in dim environments by addressing their inherent signal degradations, such as luminance attenuation and structural corruption. Although numerous algorithms attempt to improve image quality, existing methods often cause a severe loss of intrinsic signal priors. To overcome these challenges, we propose a Dual-Stream Transformer Network (DST-Net) based on illumination-agnostic signal prior guidance and multi-scale spatial convolutions. First, to address the loss of critical signal features under low-light conditions, we design a feature extraction module. This module integrates Difference of Gaussians (DoG), LAB color space transformations, and VGG-16 for texture extraction, utilizing decoupled illumination-agnostic features as signal priors to continuously guide the enhancement process. Second, we construct a dual-stream interaction architecture. By employing a cross-modal attention mechanism, the network leverages the extracted priors to dynamically rectify the deteriorated signal representation of the enhanced image, ultimately achieving iterative enhancement through differentiable curve estimation. Furthermore, to overcome the inability of existing methods to preserve fine structures and textures, we propose a Multi-Scale Spatial Fusion Block (MSFB) featuring pseudo-3D and 3D gradient operator convolutions. This module integrates explicit gradient operators to recover high-frequency edges while capturing inter-channel spatial correlations via multi-scale spatial convolutions. Extensive evaluations and ablation studies demonstrate that DST-Net achieves superior performance in subjective visual quality and objective metrics. Specifically, our method achieves a PSNR of 25.64 dB on the LOL dataset. Subsequent validation on the LSRW dataset further confirms its robust cross-scene generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。