让AI像人一样看图思考,精准找出图像质量问题并给出改进方案。
Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement

- 采用图文交替推理机制,主动获取裁剪放大等视觉证据
- 在多个数据集上超越现有方法,对生成图像质量评估更准
- 适合需要可解释性反馈的图像生成与修复场景
图像质量评估(IQA)模型正被用于指导生成模型与图像修复,这不仅要求评分准确,还需提供可操作、局部化的反馈。然而,现有基于多模态大模型的方法依赖单一语言输出,脱离人类寻证判断逻辑,导致解释性弱,难以支撑闭环优化。本文提出Q-DeepSight,一种模拟人类‘边看边想’过程的思考式框架,通过工具增强的图文链式推理(iMCoT),主动执行裁剪、放大等操作以定位质量下降区域及原因。为训练长序列推理轨迹,引入感知课程奖励(PCR)缓解奖励稀疏问题,并设计证据梯度过滤(EGF)提升视觉依据的归因准确性。Q-DeepSight在自然图像、修复图像及AI生成内容等多个基准上达到领先性能。进一步,我们构建无需训练的感知生成框架(PiG),利用Q-DeepSight的诊断结果驱动迭代优化,实现评估与修复的闭环协同。
原文摘要 · Abstract (English)
Image Quality Assessment (IQA) models are increasingly deployed as perceptual critics to guide generative models and image restoration. This role demands not only accurate scores but also actionable, localized feedback. However, current MLLM-based methods adopt a single-look, language-only paradigm, which departs from human evidence-seeking judgment and yields weakly grounded rationales, limiting their reliability for in-the-loop refinement. We propose Q-DeepSight, a think-with-image framework that emulates this human-like process. It performs interleaved Multimodal Chain-of-Thought (iMCoT) with tool-augmented evidence acquisition (e.g., crop-and-zoom) to explicitly determine where quality degrades and why. To train these long iMCoT trajectories via reinforcement learning, we introduce two techniques: Perceptual Curriculum Reward (PCR) to mitigate reward sparsity and Evidence Gradient Filtering (EGF) to improve credit assignment for visually-grounded reasoning. Q-DeepSight achieves state-of-the-art performance across diverse benchmarks, including natural, restored, and AI-generated content. Furthermore, we demonstrate its practical value with Perceptual-in-Generation (PiG), a training-free framework where Q-DeepSight's diagnoses guide iterative image enhancement, effectively closing the loop between assessment and refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。