用视觉语言模型自动识别多种图像退化,统一修复模糊、雨滴、雾霾等问题。
Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks
- 通过视觉语言模型匹配退化描述,自动选择修复策略。
- 在SOTS和Rain100L数据集上分别提升3.74dB和1.70dB。
- 首个一站式深度展开网络,适合多场景图像修复任务。
动态图像退化(如噪声、模糊、光照不一致)因传感器限制或恶劣环境而难以处理。现有深度展开网络(DUNs)虽稳定,但需人工指定每类退化的退化矩阵,适应性差。为此,我们提出视觉语言引导的展开网络(VLU-Net),一种可同时处理多种退化的统一框架。VLU-Net利用在退化图像-文本对上微调的视觉语言模型(VLM),将图像特征与退化描述对齐,自动选择对应变换。通过将基于VLM的梯度估计融入近端梯度下降(PGD)算法,有效应对复杂多退化修复任务并保持可解释性。此外,设计分层特征展开结构,高效合成不同层级的退化模式。VLU-Net是首个一站式DUN框架,在SOTS去雾数据集上优于现有方法3.74 dB,Rain100L去雨数据集上提升1.70 dB。
原文摘要 · Abstract (English)
Dynamic image degradations, including noise, blur and lighting inconsistencies, pose significant challenges in image restoration, often due to sensor limitations or adverse environmental conditions. Existing Deep Unfolding Networks (DUNs) offer stable restoration performance but require manual selection of degradation matrices for each degradation type, limiting their adaptability across diverse scenarios. To address this issue, we propose the Vision-Language-guided Unfolding Network (VLU-Net), a unified DUN framework for handling multiple degradation types simultaneously. VLU-Net leverages a Vision-Language Model (VLM) refined on degraded image-text pairs to align image features with degradation descriptions, selecting the appropriate transform for target degradation. By integrating an automatic VLM-based gradient estimation strategy into the Proximal Gradient Descent (PGD) algorithm, VLU-Net effectively tackles complex multi-degradation restoration tasks while maintaining interpretability. Furthermore, we design a hierarchical feature unfolding structure to enhance VLU-Net framework, efficiently synthesizing degradation patterns across various levels. VLU-Net is the first all-in-one DUN framework and outperforms current leading one-by-one and all-in-one end-to-end methods by 3.74 dB on the SOTS dehazing dataset and 1.70 dB on the Rain100L deraining dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。