arXiv:2609.08316cs.CV2026-09

让AI像人一样逐步发现图像质量问题的根源和程度。

From Glance to Scrutiny: Progressive Distortion Reasoning for Fine-Grained Image Quality Assessment

论文配图:From Glance to Scrutiny: Progressive Distortion Reasoning for Fine-Grained Image Quality Assessment
图 1 · 摘自论文原文
  • 分两阶段强化学习,先定位再深入分析失真
  • 在25K样本上实现精准定位、识别与严重度判断
  • 适合需要解释性图像质量评估的场景

多模态大语言模型在图像质量评估中展现出巨大潜力,但现有方法多关注整体质量预测,难以揭示失真位置与影响,限制了对局部异质退化的细粒度分析。本文提出GS-IQA框架,将IQA重构为从初步观察到深入检查的渐进式诊断过程,模拟人类感知路径。第一阶段通过感知门控奖励,先定位并识别失真区域,仅当两者均正确时才激活严重度反馈;第二阶段引入在线奖励引导的退化生成,针对模型感知瓶颈合成难例,增强对细微严重度差异的分辨能力。为系统评估,构建包含约25,000张图像的Diag-Bench基准,覆盖12种失真类型与5个等级的严重度。大量实验表明,GS-IQA在失真定位、识别与严重度估计上持续优于现有方法,且其诊断表征可有效迁移至外部全局质量预测任务。代码与数据将公开。

原文摘要 · Abstract (English)

Multi-modal large language models (MLLMs) have demonstrated significant potential in image quality assessment (IQA) by bridging visual perception with descriptive evaluations. However, existing approaches mainly focus on holistic quality prediction, often functioning as black boxes that provide limited insight into where distortions occur and how they affect perceived quality, hindering fine-grained analysis of localized and heterogeneous degradations. We propose GS-IQA, a framework that reformulates IQA as a progressive Where--What--How diagnosis, emulating the human perceptual process from an initial glance to closer scrutiny. Since a severity judgment is meaningful only for a correctly localized and recognized region, we realize this progression through a two-stage reinforcement learning paradigm that respects such dependencies: the glance stage uses a perception-gated reward to establish where degradations lie and what they are, activating severity feedback only once both are correct, while the scrutiny stage introduces online reward-conditioned degradation generation to synthesize hard examples targeted at the model's perceptual bottlenecks, sharpening its discrimination of subtle severity variations. To enable systematic evaluation, we construct Diag-Bench, a region-level IQA benchmark of about 25K curated samples spanning 12 distortion types and five ordinal severity levels. Extensive experiments show that GS-IQA consistently surpasses state-of-the-art methods in distortion localization, recognition, and severity estimation, and that its diagnostic representations transfer effectively to conventional global quality prediction across diverse external benchmarks. Code and data will be released.

图像质量评估视觉诊断强化学习可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。