arXiv:2511.05393cs.CV2025-11中稿 · EMNLP

用强化学习统一评分与排序,让AI更懂人眼对画质的判断。

PreResQ-R1: Response-Preference Disentangled Ranking-and-Scoring Reinforcement Optimization for Robust Visual Quality Assessment

  • 分离响应偏好与跨样本排序,用强化学习优化推理过程。
  • 仅用6K图像28K视频训练,10个图像/5个视频质量评估任务领先。
  • 生成可解释的推理链,揭示人眼判断画质的关键线索。

视觉质量评估(QA)旨在预测人类对视觉保真度的感知判断。尽管近期多模态大语言模型(MLLMs)在图像和视频质量推理方面展现出潜力,但现有方法主要依赖监督微调或仅排名目标,导致推理浅显、评分校准差、跨领域泛化能力弱。我们提出PreResQ-R1,一种响应-偏好解耦的强化学习框架,将绝对评分回归与相对排序一致性统一于单一推理驱动优化方案中。不同于以往方法,PreResQ-R1采用双分支奖励设计,分别建模样本内响应一致性与样本间偏好对齐,通过分组相对策略优化(GRPO)进行优化。该设计促进精细、稳定且可解释的链式思考推理。为拓展至动态视频,进一步设计全局时序与局部空间数据流策略。显著的是,仅在6K图像和28K视频上进行强化微调,PreResQ-R1在10个图像质量评估(IQA)和5个视频质量评估(VQA)基准上均达到顶尖水平,相较基线在SRCC和PLCC指标上分别提升5.30%和2.15%。除量化性能提升外,其生成的人类对齐推理轨迹揭示了质量判断背后的感知线索。

原文摘要 · Abstract (English)

Visual Quality Assessment (QA) seeks to predict human perceptual judgments of visual fidelity. While recent multimodal large language models (MLLMs) show promise in reasoning about image and video quality, existing approaches mainly rely on supervised fine-tuning or rank-only objectives, resulting in shallow reasoning, poor score calibration, and limited cross-domain generalization. We propose PreResQ-R1, a Preference-Response Disentangled Reinforcement Learning framework that unifies absolute score regression and relative ranking consistency within a single reasoning-driven optimization scheme. Unlike prior QA methods, PreResQ-R1 introduces a dual-branch reward formulation that separately models intra-sample response coherence and inter-sample preference alignment, optimized via Group Relative Policy Optimization (GRPO). This design encourages fine-grained, stable, and interpretable chain-of-thought reasoning about perceptual quality. To extend beyond static imagery, we further design a global-temporal and local-spatial data flow strategy for Video Quality Assessment. Remarkably, with reinforcement fine-tuning on only 6K images and 28K videos, PreResQ-R1 achieves state-of-the-art results across 10 IQA and 5 VQA benchmarks under both SRCC and PLCC metrics, surpassing by margins of 5.30% and textbf2.15% in IQA task, respectively. Beyond quantitative gains, it produces human-aligned reasoning traces that reveal the perceptual cues underlying quality judgments.

视觉质量评估强化学习多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。