用因果干预解决暗光图像增强中的分布偏移问题
CIVQLLIE: Causal Intervention with Vector Quantization for Low-Light Image Enhancement
- 通过向量量化构建独立于退化的视觉令牌代码本
- 多级因果干预修复输入与代码本间的分布差异
- 适合需要可解释性与真实场景鲁棒性的图像增强任务
夜间拍摄的图像因可视性严重降低而影响内容识别。现有低光图像增强(LLIE)方法面临两大挑战:数据驱动的端到端映射网络缺乏可解释性或依赖不可靠先验,在极端黑暗条件下表现不佳;基于物理的方法依赖简化假设,难以应对复杂现实场景。为此,我们提出CIVQLLIE框架,利用因果推理实现离散表示学习。通过向量量化(VQ)将连续图像特征映射到从大规模高质量图像中学习的离散代码本,该代码本编码了与退化无关的标准亮度和颜色模式。然而,退化输入与学习代码本之间的分布偏移导致直接应用失败。因此,我们提出多层级因果干预方法系统性修正这些偏移:编码阶段,像素级因果干预(PCI)模块干预以对齐低层特征与代码本期望的亮度和颜色分布;特征感知因果干预(FCI)结合低频选择性注意力门控(LSAG),识别并增强受光照退化影响最大的通道,促进准确的代码本匹配,并通过灵活的特征级干预提升编码器泛化能力;解码阶段,高频细节重建模块(HDRM)利用匹配代码本表示中保留的结构信息,通过可变形卷积技术重建精细细节。
原文摘要 · Abstract (English)
Images captured in nighttime scenes suffer from severely reduced visibility, hindering effective content perception. Current low-light image enhancement (LLIE) methods face significant challenges: data-driven end-to-end mapping networks lack interpretability or rely on unreliable prior guidance, struggling under extremely dark conditions, while physics-based methods depend on simplified assumptions that often fail in complex real-world scenarios. To address these limitations, we propose CIVQLLIE, a novel framework that leverages the power of discrete representation learning through causal reasoning. We achieve this through Vector Quantization (VQ), which maps continuous image features to a discrete codebook of visual tokens learned from large-scale high-quality images. This codebook serves as a reliable prior, encoding standardized brightness and color patterns that are independent of degradation. However, direct application of VQ to low-light images fails due to distribution shifts between degraded inputs and the learned codebook. Therefore, we propose a multi-level causal intervention approach to systematically correct these shifts. First, during encoding, our Pixel-level Causal Intervention (PCI) module intervenes to align low-level features with the brightness and color distributions expected by the codebook. Second, a Feature-aware Causal Intervention (FCI) mechanism with Low-frequency Selective Attention Gating (LSAG) identifies and enhances channels most affected by illumination degradation, facilitating accurate codebook token matching while enhancing the encoder's generalization performance through flexible feature-level intervention. Finally, during decoding, the High-frequency Detail Reconstruction Module (HDRM) leverages structural information preserved in the matched codebook representations to reconstruct fine details using deformable convolution techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。