首个融合推理、定位与指代的多粒度图像质量评估框架
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring

- 构建四任务统一范式,覆盖全局/局部描述、像素级定位与区域指代
- 提出无冲突两阶段训练方案,实现从文本到像素的渐进式对齐
- 支持可解释性评估,适用于需要精细质量分析的场景
我们提出IQA-Spider,首个将推理、定位与指代统一于单个基于视觉语言模型(LMM)框架中的多粒度图像质量评估方法。现有基于LMM的IQA方法通常仅支持部分感知维度,如质量描述或问答(即推理),或像素级定位,其局限性主要源于缺乏统一的任务与数据范式,以及多粒度学习的有效优化策略。为此,我们设计了一个涵盖全局与局部质量描述、像素级定位、区域级指代的严谨四任务范式,并构建了具有可扩展自动标注流程的IQA数据集,为统一多粒度学习提供坚实基础。为进一步实现统一感知,采用无冲突的两阶段设计:第一阶段通过多任务学习提升模型在细粒度文本层面的理解能力;第二阶段引入无需训练的文本到点定位机制,通过将标记概率映射至空间坐标,建立文本语义与像素感知间的桥梁。基于上述工作,IQA-Spider实现了统一的可解释多粒度图像质量评估。在多个基准上的大量实验验证了该范式与框架的有效性与通用性。
原文摘要 · Abstract (English)
We present IQA-Spider, the first image quality assessment (IQA) framework that unifies reasoning, grounding, and referring into a single LMM-based framework for multi-granularity quality understanding. Existing LMM-based IQA methods typically support only partial perception dimensions, such as quality description and question answering~(\textit{i.e.}, reasoning) or pixel-level grounding. This limitation largely stems from the absence of (i) a unified task and data formulation and (ii) effective optimization paradigms for multi-granularity learning. To address these limitations, we formulate a rigorous four-task paradigm covering global and local quality description, pixel-level grounding, and region-level referring. Based on this formulation, we construct a corresponding IQA dataset with a scalable and automatic annotation pipeline, thereby providing a solid foundation for unified multi-granularity learning. To further enable unified perception, we adopt a conflict-free two-stage design that progressively extends text-level multi-granularity understanding to pixel-level grounding: (i) the first stage equips the model with fine-grained text-level reasoning across multiple IQA tasks, and (ii) the second stage introduces a training-free text-to-point grounding paradigm, which bridges textual semantics and pixel-level perception by mapping token logits to spatial coordinates. Based on these efforts, we achieve IQA-Spider with unified multi-granularity explainable image quality assessment. Extensive experiments across multiple benchmarks demonstrate strong performance, validating the effectiveness and versatility of the proposed formulation and framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。