arXiv:2605.12556cs.CV2026-05中稿 · 2026 IEEE Internat…被引 1

融合深度、亮度和语义信息,提升暗光图像增强效果

M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement

论文配图:M2Retinexformer: Multi-Modal Retinexformer for Low-Light Image Enhancement
图 1 · 摘自论文原文
  • 多模态信息通过交叉注意力融合,动态调节光照引导
  • 在LOL等4个数据集上超越现有方法,细节更清晰噪声更少
  • 适合图像增强研究者和需要高画质低光处理的工程师

暗光图像增强因噪声放大、伪影和色彩失真等复杂退化而具有挑战性。尽管基于Retinex的深度学习方法已取得良好效果,但主要依赖单一的RGB信息。本文提出M2Retinexformer框架,通过在渐进式优化流程中引入深度图、亮度先验和语义特征,扩展了Retinexformer。深度提供不受光照变化影响的几何上下文,亮度与语义特征则分别指导亮度分布与场景理解。多尺度提取各模态信息,通过交叉注意力融合,并采用自适应门控机制,根据辅助线索可靠性动态平衡光照引导的自注意力与跨模态注意力。在LOL、SID、SMID和SDSD四个基准上的评估显示,该方法整体优于Retinexformer及近期先进方法。代码与预训练权重已开源。

原文摘要 · Abstract (English)

Low-light image enhancement is challenging due to complex degradations, including amplified noise, artifacts, and color distortion. While Retinex-based deep learning methods have achieved promising results, they primarily rely on single-modality RGB information. We propose M2Retinexformer (Multi-Modal Retinexformer), a novel framework that extends Retinexformer by incorporating depth cues, luminance priors, and semantic features within a progressive refinement pipeline. Depth provides geometric context that is invariant to lighting variations, while luminance and semantic features offer explicit guidance on brightness distribution and scene understanding. Modalities are extracted at multiple scales and fused through cross-attention, with adaptive gating dynamically balancing illumination-guided self-attention and cross-attention based on the reliability of auxiliary cues. Evaluations on the LOL, SID, SMID, and SDSD benchmarks demonstrate overall improvements over Retinexformer and recent state-of-the-art methods. Code and pretrained weights are available at https://github.com/YoussefAboelwafa/M2Retinexformer

图像增强多模态暗光处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。