提出多粒度图像质量评估框架,同时判断整体质量和细节属性。
Multi-Granularity Reasoning for Image Quality Assessment via Attribute-Aware Reinforcement Learning to Rank
- 用属性感知提示引导模型生成分维度推理。
- 在8个基准上整体评分提升2.1%(SRCC),属性评估更准确。
- 适合需要可解释性质量评估的AI生成内容审核场景。
近期基于推理的图像质量评估(IQA)研究展示了强化学习排序(RL2R)在训练视觉语言模型(VLMs)进行感知质量评估方面的潜力。然而,现有方法仅在单一粒度下工作,仅预测整体质量分数,忽略了人类质量感知的多维特性,如清晰度、色彩保真度、噪声水平和构图美学等。本文提出MG-IQA(多粒度IQA),一种将RL2R扩展至单次推理中联合评估整体质量与细粒度质量属性的多粒度推理框架。本方法引入三项关键创新:(1)属性感知提示策略,激发VLMs产生结构化多属性推理;(2)多维Thurstone奖励模型,为群体相对策略优化计算属性特异性保真度奖励;(3)跨域对齐机制,使合成失真、真实失真及AI生成图像数据集间无需感知尺度重校准即可稳定联合训练。在八个IQA基准上的大量实验表明,MG-IQA在整体质量预测(平均SRCC提升2.1%)和属性级评估上均持续优于现有最佳方法,同时生成可解释且符合人类认知的质量描述。
原文摘要 · Abstract (English)
Recent advances in reasoning-induced image quality assessment (IQA) have demonstrated the power of reinforcement learning to rank (RL2R) for training vision-language models (VLMs) to assess perceptual quality. However, existing approaches operate at a single granularity, predicting only an overall quality score, while overlooking the multi-dimensional nature of human quality perception, which encompasses attributes such as sharpness, color fidelity, noise level, and compositional aesthetics. In this paper, we propose MG-IQA (Multi-Granularity IQA), a multi-granularity reasoning framework that extends RL2R to jointly assess overall image quality and fine-grained quality attributes within a single inference pass. Our approach introduces three key innovations: (1) an attribute-aware prompting strategy that elicits structured multi-attribute reasoning from VLMs; (2) a multi-dimensional Thurstone reward model that computes attribute-specific fidelity rewards for group relative policy optimization; and (3) a cross-domain alignment mechanism that enables stable joint training across synthetic distortion, authentic distortion, and AI-generated image datasets without perceptual scale re-alignment. Extensive experiments on eight IQA benchmarks demonstrate that MG-IQA consistently outperforms state-of-the-art methods in both overall quality prediction (average SRCC improvement of 2.1\%) and attribute-level assessment, while generating interpretable, human-aligned quality descriptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。