arXiv:2504.07148eess.IV2025-04被引 13

用思维链提升图像修复质量,自动识别并最优排序修复步骤。

Q-Agent: Quality-Driven Chain-of-Thought Image Restoration Agent through Robust Multimodal Large Language Model

  • 通过思维链分解多退化感知,增强大模型对复杂退化的理解。
  • 结合客观质量评估,智能选择修复顺序,性能优于现有全功能模型。
  • 适合需要高效精准修复的现实场景,如医疗影像、老旧照片恢复。

真实世界中的图像修复常面临噪声、模糊、压缩伪影和低分辨率等多种未知退化问题。针对特定退化训练专用模型会导致泛化能力差;而全功能模型虽能处理多种退化,却在某些类型上性能下降,且难以应对训练中未见的退化。现有基于多模态大语言模型(MLLM)的修复代理依赖耗时的回溯选择策略,忽视图像质量,易误判退化,并产生冗余操作与高计算开销。为此,我们提出质量驱动型修复代理(Q-Agent),采用思维链(CoT)修复框架。其包含鲁棒退化感知与质量驱动的贪心修复两部分:前者微调MLLM,并利用CoT将多退化感知拆解为单退化任务以增强感知能力;后者通过客观图像质量评估(IQA)指标确定最优修复顺序并执行对应算法。实验表明,Q-Agent在多项指标上显著优于现有全功能模型。

原文摘要 · Abstract (English)

Image restoration (IR) often faces various complex and unknown degradations in real-world scenarios, such as noise, blurring, compression artifacts, and low resolution, etc. Training specific models for specific degradation may lead to poor generalization. To handle multiple degradations simultaneously, All-in-One models might sacrifice performance on certain types of degradation and still struggle with unseen degradations during training. Existing IR agents rely on multimodal large language models (MLLM) and a time-consuming rolling-back selection strategy neglecting image quality. As a result, they may misinterpret degradations and have high time and computational costs to conduct unnecessary IR tasks with redundant order. To address these, we propose a Quality-Driven agent (Q-Agent) via Chain-of-Thought (CoT) restoration. Specifically, our Q-Agent consists of robust degradation perception and quality-driven greedy restoration. The former module first fine-tunes MLLM, and uses CoT to decompose multi-degradation perception into single-degradation perception tasks to enhance the perception of MLLMs. The latter employs objective image quality assessment (IQA) metrics to determine the optimal restoration sequence and execute the corresponding restoration algorithms. Experimental results demonstrate that our Q-Agent achieves superior IR performance compared to existing All-in-One models.

图像修复思维链多模态大模型质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。