提出新数据集与模型,显著提升真实场景低光图像增强效果。
Illuminating Darkness: Learning to Enhance Low-light Images In-the-Wild
- 分离亮度与色度编码,减少颜色与结构混淆
- 在LSD数据集上提升2.45 dB PSNR,检测任务提升6.80 mAP
- 适合低光图像处理、手机摄影优化等应用场景
单帧低光图像增强因真实世界成对数据稀缺仍具挑战。为此,我们构建了大规模野外采集的低光智能手机数据集(LSD),覆盖0.1至200勒克斯多种光照条件,包含6,425对精确配准的低光与正常光照图像,源自超过8,000个动态室内外场景,通过多帧采集与专家评估筛选。为评估泛化与美学质量,另收集2,117张未见设备的无配对低光图像。为充分利用LSD,提出TFFormer:通过分离亮度与色度(LC)编码降低耦合,设计交叉注意力驱动的联合解码器实现上下文感知融合,并引入LC精炼与LC引导监督,显著提升感知保真度与结构一致性。TFFormer在LSD上达最新最优结果(+2.45 dB PSNR),并大幅改善下游视觉任务,如低光目标检测在ExDark上提升6.80 mAP。
原文摘要 · Abstract (English)
Single-shot low-light image enhancement (SLLIE) remains challenging due to the limited availability of diverse, real-world paired datasets. To bridge this gap, we introduce the Low-Light Smartphone Dataset (LSD), a large-scale, high-resolution (4K+) dataset collected in the wild across a wide range of challenging lighting conditions (0.1 to 200 lux). LSD contains 6,425 precisely aligned low and normal-light image pairs, selected from over 8,000 dynamic indoor and outdoor scenes through multi-frame acquisition and expert evaluation. To evaluate generalization and aesthetic quality, we collect 2,117 unpaired low-light images from previously unseen devices. To fully exploit LSD, we propose TFFormer, a hybrid model that encodes luminance and chrominance (LC) separately to reduce color-structure entanglement. We further propose a cross-attention-driven joint decoder for context-aware fusion of LC representations, along with LC refinement and LC-guided supervision to significantly enhance perceptual fidelity and structural consistency. TFFormer achieves state-of-the-art results on LSD (+2.45 dB PSNR) and substantially improves downstream vision tasks, such as low-light object detection (+6.80 mAP on ExDark).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。