用视觉语言模型自动识别图像退化类型并精准修复
Degradation-Aware Image Enhancement via Vision-Language Classification
- 通过VLM将图像分为四类退化类型
- 针对不同退化使用专用模型修复,提升视觉质量
- 适合需要自动化图像增强的工程场景
图像退化在各类实际应用中普遍存在,影响视觉质量和下游任务。本文提出一种新框架,利用视觉语言模型(VLM)自动将输入图像分类为四种预定义退化类型:(A) 超分辨率退化(含噪声、模糊和JPEG压缩)、(B) 反射伪影、(C) 运动模糊或 (D) 无明显退化(高质量图像)。分类后,属于类别A、B或C的图像分别由对应专用模型进行针对性恢复,最终输出视觉质量更优的图像。实验表明,该方法能准确分类图像退化类型,并通过专用恢复模型显著提升图像质量。本方法为真实世界图像增强提供了可扩展、自动化的解决方案,结合了VLM的能力与先进恢复技术。
原文摘要 · Abstract (English)
Image degradation is a prevalent issue in various real-world applications, affecting visual quality and downstream processing tasks. In this study, we propose a novel framework that employs a Vision-Language Model (VLM) to automatically classify degraded images into predefined categories. The VLM categorizes an input image into one of four degradation types: (A) super-resolution degradation (including noise, blur, and JPEG compression), (B) reflection artifacts, (C) motion blur, or (D) no visible degradation (high-quality image). Once classified, images assigned to categories A, B, or C undergo targeted restoration using dedicated models tailored for each specific degradation type. The final output is a restored image with improved visual quality. Experimental results demonstrate the effectiveness of our approach in accurately classifying image degradations and enhancing image quality through specialized restoration models. Our method presents a scalable and automated solution for real-world image enhancement tasks, leveraging the capabilities of VLMs in conjunction with state-of-the-art restoration techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。