arXiv:2510.10609cs.CV2025-10

统一视觉质量评估,让模型自动给出可解释的评分

OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment

  • 构建多维度推理链数据集,支持连续可解释评分
  • 在三大图像质量任务上均优于现有方法
  • 适合需要高质量评分的生成模型优化场景

当前视觉评估方法通常局限于单一任务。为此,我们提出OmniQuality-R,一个统一的奖励建模框架,将多任务质量判断转化为连续且可解释的奖励信号,用于策略优化。受主观实验启发,参与者在评估前会获得特定任务的评价指引。我们据此设计结构化奖励建模框架,将多维推理转化为连续可解释的奖励信号。为实现此目标,我们通过拒绝采样抽取具有信息量的计划-推理轨迹,构建增强型推理奖励数据集,用于监督微调(SFT)。在此基础上,采用基于高斯分布的组相对策略优化(GRPO)进行后训练,支持连续分数预测。为进一步稳定训练并提升下游泛化能力,我们在强化学习中引入标准差过滤和熵门控机制,抑制不稳定的更新,降低策略优化方差。我们在三个关键图像质量评估任务上评估OmniQuality-R:审美质量评估、技术质量评价和文本-图像对齐。

原文摘要 · Abstract (English)

Current visual evaluation approaches are typically constrained to a single task. To address this, we propose OmniQuality-R, a unified reward modeling framework that transforms multi-task quality reasoning into continuous and interpretable reward signals for policy optimization. Inspired by subjective experiments, where participants are given task-specific instructions outlining distinct assessment principles prior to evaluation, we propose OmniQuality-R, a structured reward modeling framework that transforms multi-dimensional reasoning into continuous and interpretable reward signals. To enable this, we construct a reasoning-enhanced reward modeling dataset by sampling informative plan-reason trajectories via rejection sampling, forming a reliable chain-of-thought (CoT) dataset for supervised fine-tuning (SFT). Building on this, we apply Group Relative Policy Optimization (GRPO) for post-training, using a Gaussian-based reward to support continuous score prediction. To further stabilize the training and improve downstream generalization, we incorporate standard deviation (STD) filtering and entropy gating mechanisms during reinforcement learning. These techniques suppress unstable updates and reduce variance in policy optimization. We evaluate OmniQuality-R on three key IQA tasks: aesthetic quality assessment, technical quality evaluation, and text-image alignment.

奖励模型图像质量强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。