arXiv:2512.06750cs.CV2025-12被引 5

首个统一图像质量评估与修复的视觉语言模型,让评分指导修复。

UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement

  • 两阶段训练:从单一退化到复杂混合退化逐步提升鲁棒性。
  • 多任务联合训练使修复性能显著提升,尤其在复杂退化场景下。
  • 适合需要高质量图像修复与评估的一体化应用开发人员。

图像质量评估(IQA)与图像修复是低层视觉的基础问题。尽管两者概念紧密相关,但现有方法通常将其孤立处理。近期统一多模态理解-生成模型的进步表明,更强的理解能力可提升生成效果。这促使我们探索一个统一的模型,同时实现IQA与修复,并明确研究如何用IQA引导修复——这一方向虽潜力巨大却尚未充分挖掘。本文提出UARE,据我们所知首个统一图像质量评估、修复与增强的视觉语言模型。基于预训练的统一理解与生成模型,设计两阶段训练框架:第一阶段采用由易到难的渐进式策略,从单一退化类型扩展至高阶混合退化,使UARE能应对多种退化;第二阶段通过交错文本-图像数据进行统一微调,对齐IQA信号与修复目标。通过多任务协同训练,UARE利用质量评估信息显著提升修复与增强性能。在IQA、修复与增强任务上的大量实验验证了其有效性。代码与模型将公开于https://github.com/lwq20020127/UARE。

原文摘要 · Abstract (English)

Image quality assessment (IQA) and image restoration are fundamental problems in low-level vision. Although IQA and restoration are closely connected conceptually, most existing work treats them in isolation. Recent advances in unified multimodal understanding-generation models demonstrate promising results and indicate that stronger understanding can improve generative performance. This motivates a single model that unifies IQA and restoration and explicitly studies how IQA can guide restoration, a setting that remains largely underexplored yet highly valuable. In this paper, we propose UARE, to our knowledge the first Unified vision-language model for image quality Assessment, Restoration, and Enhancement. Built on pretrained unified understanding and generation models, we introduce a two-stage training framework. First, a progressive, easy-to-hard schedule expands from single-type distortions to higher-order mixed degradations, enabling UARE to handle multiple degradations. Second, we perform unified fine-tuning of quality understanding and restoration with interleaved text-image data, aligning IQA signals with restoration objectives. Through multi-task co-training, UARE leverages IQA to boost restoration and enhancement performance. Extensive experiments across IQA, restoration, and enhancement tasks demonstrate the effectiveness of UARE. The code and models will be available at https://github.com/lwq20020127/UARE.

图像修复多任务学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。