无需配对数据,一键修复暗光模糊图像。
Zero-Reference Joint Low-Light Enhancement and Deblurring via Visual Autoregressive Modeling with VLM-Derived Modulation
- 用视觉语言模型生成光照提示,指导图像修复
- 在多个基准上超越现有方法,暗光模糊图像修复效果最优
- 适合无标注数据场景,尤其适合真实拍摄的低质图像
真实世界中的暗光图像常伴随低亮度、低对比度、复杂噪声与运动模糊,修复难度大。现有方法多依赖成对训练数据或无法建模动态光照与模糊特征,泛化能力差。本文提出一种基于视觉自回归(VAR)建模的生成框架,利用视觉语言模型(VLM)提供的感知先验作为引导。为向VAR提供有效条件信号,设计自适应曲线估计策略,根据VLM生成的可见性评分调节多样光照。同时,将动态空间频率感知的旋转位置编码(SF-RoPE)融入VAR,增强对模糊退化的结构建模能力。此外,提出递归相域调制策略,通过受VLM评估模糊度约束的有界迭代优化,减轻相域中的模糊伪影。整个框架完全无监督,在多个基准数据集上达到当前最佳性能。
原文摘要 · Abstract (English)
Real-world dark images commonly exhibit not only low visibility and contrast but also complex noise and blur, posing significant restoration challenges. Existing methods often rely on paired data or fail to model dynamic illumination and blur characteristics, leading to poor generalization. To tackle this, we propose a generative framework based on visual autoregressive (VAR) modeling, guided by perceptual priors from the vision-language model (VLM). Specifically, to supply informative conditioning cues for VAR models, we deploy an adaptive curve estimation scheme to modulate the diverse illumination based on VLM-derived visibility scores. In addition, we integrate dynamic and spatial-frequency-aware Rotary Positional Encodings (SF-RoPE) into VAR to enhance its ability to model structures degraded by blur. Furthermore, we propose a recursive phase-domain modulation strategy that mitigates blur-induced artifacts in the phase domain via bounded iterative refinement guided by VLM-assessed blur scores. Our framework is fully unsupervised and achieves state-of-the-art performance on benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。