arXiv:2504.15054cs.CV2025-04中稿 · IEEE Transactions …被引 10

用结构引导的扩散Transformer提升暗光图像质量,降噪同时增强细节。

Structure-guided Diffusion Transformer for Low-Light Image Enhancement

  • 通过小波变换压缩特征,提升推理效率并捕捉多向频带。
  • 引入结构增强模块与自适应融合,精准恢复纹理细节。
  • 设计结构引导注意力块,抑制噪声区域干扰,适合暗光图像处理场景。

尽管扩散Transformer(DiT)近年来备受关注,但在暗光图像增强中的应用仍属空白。现有方法在恢复细节的同时不可避免放大噪声,导致视觉质量下降。本文首次将DiT引入暗光增强任务,提出结构引导的扩散Transformer框架(SDTL)。通过小波变换压缩特征,提升模型推理效率并捕捉多方向频带信息;设计结构增强模块(SEM),利用结构先验增强纹理,并采用自适应融合策略实现更精确的增强效果;此外,提出结构引导注意力块(SAB),聚焦纹理丰富特征,避免噪声区域干扰。大量定性与定量实验表明,该方法在多个主流数据集上达到最优性能,验证了SDTL在提升图像质量方面的有效性,以及DiT在暗光增强任务中的潜力。

原文摘要 · Abstract (English)

While the diffusion transformer (DiT) has become a focal point of interest in recent years, its application in low-light image enhancement remains a blank area for exploration. Current methods recover the details from low-light images while inevitably amplifying the noise in images, resulting in poor visual quality. In this paper, we firstly introduce DiT into the low-light enhancement task and design a novel Structure-guided Diffusion Transformer based Low-light image enhancement (SDTL) framework. We compress the feature through wavelet transform to improve the inference efficiency of the model and capture the multi-directional frequency band. Then we propose a Structure Enhancement Module (SEM) that uses structural prior to enhance the texture and leverages an adaptive fusion strategy to achieve more accurate enhancement effect. In Addition, we propose a Structure-guided Attention Block (SAB) to pay more attention to texture-riched tokens and avoid interference from noisy areas in noise prediction. Extensive qualitative and quantitative experiments demonstrate that our method achieves SOTA performance on several popular datasets, validating the effectiveness of SDTL in improving image quality and the potential of DiT in low-light enhancement tasks.

图像增强扩散模型低光照注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。