arXiv:2504.08219cs.CV2025-04被引 4

用视觉语言模型统一修复各种恶劣天气下的图像退化

VL-UR: Vision-Language-guided Universal Restoration of Images Degraded by Adverse Weather Conditions

  • 通过CLIP模型融合图像与语义信息,实现跨场景自适应修复
  • 在11种退化条件下达到最优性能,显著提升鲁棒性
  • 适合自动驾驶、安防监控等真实复杂环境应用

图像恢复对提升退化图像质量至关重要,广泛应用于自动驾驶、安全监控和数字内容增强等领域。然而,现有方法多针对特定退化场景,难以应对真实环境中多样且复杂的退化问题。此外,真实退化的非均匀性凸显了对自适应智能解决方案的需求。为此,我们提出一种新的视觉-语言引导通用修复框架(VL-UR)。VL-UR利用零样本对比语言图像预训练(CLIP)模型,通过整合视觉与语义信息提升图像恢复效果。引入场景分类器以适配CLIP,生成与退化图像对齐的高质量语言嵌入,并预测复杂场景中的退化类型。在11种不同退化设置下的大量实验表明,VL-UR在性能、鲁棒性和适应性方面均达到当前最优水平,为动态真实环境中的现代图像恢复挑战提供了变革性解决方案。

原文摘要 · Abstract (English)

Image restoration is critical for improving the quality of degraded images, which is vital for applications like autonomous driving, security surveillance, and digital content enhancement. However, existing methods are often tailored to specific degradation scenarios, limiting their adaptability to the diverse and complex challenges in real-world environments. Moreover, real-world degradations are typically non-uniform, highlighting the need for adaptive and intelligent solutions. To address these issues, we propose a novel vision-language-guided universal restoration (VL-UR) framework. VL-UR leverages a zero-shot contrastive language-image pre-training (CLIP) model to enhance image restoration by integrating visual and semantic information. A scene classifier is introduced to adapt CLIP, generating high-quality language embeddings aligned with degraded images while predicting degraded types for complex scenarios. Extensive experiments across eleven diverse degradation settings demonstrate VL-UR's state-of-the-art performance, robustness, and adaptability. This positions VL-UR as a transformative solution for modern image restoration challenges in dynamic, real-world environments.

图像修复视觉语言通用恢复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。