arXiv:2508.03763cs.CVcs.AI2025-08AAAI被引 7

通过多阶段强化微调,提升图像质量评估的感知与推理能力

Refine-IQA: Multi-Stage Reinforcement Finetuning for Perceptual Image Quality Assessment

  • 构建2万张带局部失真图像的数据集,多任务奖励增强视觉感知
  • 引入概率差异奖励,监督模型思考过程,提升评分准确性
  • 首次实现高质量图像解释能力,适合需要可解释性评估的研究者

强化微调(RFT)是大模型训练的新兴范式。类似高层次推理任务,其也可应用于低层视觉领域,如图像质量评估(IQA)。现有基于RFT的IQA方法通常使用规则生成输出奖励来验证模型推演结果,但缺乏对“思考”过程的奖励监督,导致其正确性与有效性无法控制。此外,这些方法直接在下游IQA任务上微调,未显式增强模型的原始低层视觉质量感知能力,可能限制性能上限。针对上述问题,我们提出多阶段RFT图像质量评估框架(Refine-IQA)。在第一阶段,构建了包含12种主要失真、20,907张局部失真图像和超过55,000个RFT样本的Refine-Perception-20K数据集,并设计多任务奖励函数以强化模型的视觉质量感知能力。在第二阶段,面向质量评分任务,引入概率差异奖励策略,实现对“思考”过程的有效监督。所提出的Refine-IQA系列模型在感知与评分任务上均表现优异,尤其其范式激活了稳健的“思考”(质量解释)能力,在对应的质量解释基准测试中也取得卓越成绩。

原文摘要 · Abstract (English)

Reinforcement fine-tuning (RFT) is a proliferating paradigm for LMM training. Analogous to high-level reasoning tasks, RFT is similarly applicable to low-level vision domains, including image quality assessment (IQA). Existing RFT-based IQA methods typically use rule-based output rewards to verify the model's rollouts but provide no reward supervision for the "think" process, leaving its correctness and efficacy uncontrolled. Furthermore, these methods typically fine-tune directly on downstream IQA tasks without explicitly enhancing the model's native low-level visual quality perception, which may constrain its performance upper bound. In response to these gaps, we propose the multi-stage RFT IQA framework (Refine-IQA). In Stage-1, we build the Refine-Perception-20K dataset (with 12 main distortions, 20,907 locally-distorted images, and over 55K RFT samples) and design multi-task reward functions to strengthen the model's visual quality perception. In Stage-2, targeting the quality scoring task, we introduce a probability difference reward involved strategy for "think" process supervision. The resulting Refine-IQA Series Models achieve outstanding performance on both perception and scoring tasks-and, notably, our paradigm activates a robust "think" (quality interpreting) capability that also attains exceptional results on the corresponding quality interpreting benchmark.

图像质量评估强化学习多阶段训练可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。