提升暗光图像增强的细节保真度,让恢复画面更真实清晰
Boosting Fidelity for Pre-Trained-Diffusion-Based Low-Light Image Enhancement via Condition Refinement
- 通过重构编码损失的潜在特征,增强条件信息
- 动态交互条件与噪声潜在变量,显著提升恢复精度
- 无需重训练,可直接接入现有扩散模型使用
基于预训练扩散模型(如 Stable Diffusion)的方法在低层次视觉任务中表现优异。然而,预训练扩散基(PTDB)方法常因追求感知真实感而牺牲内容保真度,尤其在暗光场景下,因光照严重退化导致控制能力下降。我们发现保真度损失主要源于缺乏合适的条件潜在建模,以及条件潜在与噪声潜在之间缺乏双向交互。为此,我们提出一种新的条件优化策略:引入潜空间修复机制,利用生成先验恢复 VAE 编码过程中丢失的空间细节;同时,使精炼后的条件潜变量与噪声潜变量动态交互,实现更精准的图像重建。该方法为即插即用设计,可无缝集成至现有扩散网络中,显著提升保真度。大量实验表明,该方法在多个基准上均取得明显性能提升。
原文摘要 · Abstract (English)
Diffusion-based methods, leveraging pre-trained large models like Stable Diffusion via ControlNet, have achieved remarkable performance in several low-level vision tasks. However, Pre-Trained Diffusion-Based (PTDB) methods often sacrifice content fidelity to attain higher perceptual realism. This issue is exacerbated in low-light scenarios, where severely degraded information caused by the darkness limits effective control. We identify two primary causes of fidelity loss: the absence of suitable conditional latent modeling and the lack of bidirectional interaction between the conditional latent and noisy latent in the diffusion process. To address this, we propose a novel optimization strategy for conditioning in pre-trained diffusion models, enhancing fidelity while preserving realism and aesthetics. Our method introduces a mechanism to recover spatial details lost during VAE encoding, i.e., a latent refinement pipeline incorporating generative priors. Additionally, the refined latent condition interacts dynamically with the noisy latent, leading to improved restoration performance. Our approach is plug-and-play, seamlessly integrating into existing diffusion networks to provide more effective control. Extensive experiments demonstrate significant fidelity improvements in PTDB methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。