基于深度引导的多阶段低光增强,提升细节与真实感
LUMEN: Low-light Unified Multi-stage Enhancement Network using depth-guided flash, clustering, and attention-based Transformers

- 先估深度再分区域,实现按深度模拟补光
- 融合注意力机制,同时增强全局上下文与局部细节
- 在LOL数据集上超越现有方法,结果更自然
低光图像增强因严重噪声、色彩失真、对比度下降和结构细节丢失而仍具挑战。现有方法通常采用统一增强策略,未考虑真实场景中光衰减与传感器噪声的深度依赖性。为此,我们提出LUMEN,一种结合虚拟闪光模拟与基于Transformer的特征融合的多阶段增强框架。该框架首先通过专用编码器-解码器网络从低光输入中估计场景深度,随后利用软聚类模块将像素划分为深度感知区域,实现深度相关的闪光模拟。模拟闪光特征与深度表示通过高效的注意力融合块与图像特征融合,以增强全局上下文并保留精细细节。复合损失函数包含重建、感知、结构、色彩、边缘和深度一致性目标,确保视觉保真度与感知质量。在LOL-v1和LOL-v2基准上的大量实验表明,LUMEN达到最优性能,生成结果在视觉上更为自然。
原文摘要 · Abstract (English)
Low-light image enhancement remains a challenging problem due to severe noise, color distortion, contrast degradation, and loss of structural details under insufficient illumination. Existing methods typically apply uniform enhancement without considering the depth-dependent nature of light attenuation and sensor noise in real-world scenes. To address this limitation, we propose LUMEN, a multi-stage enhancement framework that integrates virtual flash simulation with transformer-based feature fusion. The proposed framework first estimates scene depth from low-light inputs using a dedicated encoder-decoder network, after which a soft clustering module partitions pixels into depth-aware regions, enabling depth-dependent flash simulation. The simulated flash features, together with depth representations, are fused with image features through efficient attention-based fusion blocks to enhance global context while preserving fine details. A composite loss function combining reconstruction, perceptual, structural, color, edge, and depth consistency objectives ensures both visual fidelity and perceptual quality. Extensive experiments on LOL-v1 and LOL-v2 benchmarks demonstrate that LUMEN achieves state-of-the-art performance and produces visually natural results compared with several state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。