用视频对比提升画质评估的跨域通用性
VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment

- 通过直接比较两段视频的优劣,避免评分偏差
- 在多个基准上达到最佳性能,跨领域泛化能力强
- 适合需要可靠排序和跨场景评估的研究者
大型多模态模型在视频质量评估中展现潜力,但多数方法仍为每段视频预测单一分数。这种点式监督常混入感知质量与数据集特定校准(如标注规范、打分习惯、分数分布),导致模型在基准内表现良好但在未见领域迁移能力差。我们提出 extbf{VersusQ},一种完全基于直接对比的成对边际推理框架。该方法利用多模态模型比较两段视频,分析其视觉与时间质量差异,输出带符号的连续边际值,同时体现偏好选择与差异程度。为进一步对齐可解释的推理理由与细粒度数值差异,引入边际耦合的GRPO优化策略,联合优化基于推演的关系推理与连续边际回归。在多个公开视频质量评估基准上的实验表明,VersusQ实现最先进性能,具备强跨域泛化能力,并在异构评估场景下保持可靠细粒度排序。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have shown promise for video quality assessment, but most methods still predict an absolute score for each video. Such pointwise supervision often mixes perceptual quality with dataset-specific calibration, including annotation protocols, rating habits, and score distributions. As a result, the learned scoring rule may work well within a benchmark but transfer poorly across unseen domains. We argue that relative comparisons alleviate the absolute-scale calibration bias by focusing purely on perceptual differences rather than dataset-specific rating habits. Consequently, we propose \textbf{VersusQ}, a pairwise margin reasoning framework driven entirely by direct comparisons. Specifically, VersusQ performs LMM-based comparison between two videos, reasons about their visual and temporal quality differences, and predicts a signed continuous margin that captures both the preferred choice and the degree of difference. Furthermore, to align interpretable comparison rationales with fine-grained numerical differences, we introduce Margin-Coupled GRPO, which jointly optimizes rollout-based relational reasoning and continuous margin regression. Extensive experiments on multiple public VQA benchmarks demonstrate that VersusQ achieves state-of-the-art performance, strong cross-domain generalization, and reliable fine-grained ranking under heterogeneous evaluation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。