arXiv:2512.02224cs.CV2025-12

统一视频质量评估框架,能诊断多种画质问题并给出解释。

Towards Unified Video Quality Assessment

  • 将视频质量评估建模为专家混合模型,每个专家专注不同感知领域。
  • 在17个数据库上超越18种基准方法,对多种画质退化类型检测准确率高。
  • 适合需要可解释性反馈的视频质量分析场景,如流媒体优化与评测。

当前视频质量评估(VQA)多采用单一模型,仅输出一个质量分,无法提供可解释的诊断信息。多数方法还局限于特定格式,缺乏通用性。为此,本文提出Unified-VQA,将通用VQA重构为诊断型混合专家(MoE)问题。该框架包含多个专注于不同感知领域的「专家」,设计了一种新型多代理专家训练策略,利用各领域最适配的代理指标进行排名式优化。同时引入诊断多任务头,生成全局质量分数与可解释的多维失真向量,通过弱监督学习策略优化,利用本研究构建的大规模训练数据集的已知属性。在不重新训练或微调的情况下,Unified-VQA在17个包含高清、超清、高动态范围和高帧率等多种格式中多样流媒体失真问题的数据集上,对通用VQA和诊断性失真检测任务均表现出一致且优越的性能,显著优于18种基准方法。该工作为实现实用、可操作且可解释的视频质量评估迈出关键一步。

原文摘要 · Abstract (English)

Recent works in video quality assessment (VQA) typically employ monolithic models that typically predict a single quality score for each test video. These approaches cannot provide diagnostic, interpretable feedback, offering little insight into why the video quality is degraded. Most of them are also specialized, format-specific metrics rather than truly ``generic" solutions, as they are designed to learn a compromised representation from disparate perceptual domains. To address these limitations, this paper proposes Unified-VQA, a framework that provides a single, unified quality model applicable to various distortion types within multiple video formats by recasting generic VQA as a Diagnostic Mixture-of-Experts (MoE) problem. Unified-VQA employs multiple ``perceptual experts'' dedicated to distinct perceptual domains. A novel multi-proxy expert training strategy is designed to optimize each expert using a ranking-inspired loss, guided by the most suitable proxy metric for its domain. We also integrated a diagnostic multi-task head into this framework to generate a global quality score and an interpretable multi-dimensional artifact vector, which is optimized using a weakly-supervised learning strategy, leveraging the known properties of the large-scale training database generated for this work. With static model parameters (without retraining or fine-tuning), Unified-VQA demonstrates consistent and superior performance compared to over 18 benchmark methods for both generic VQA and diagnostic artifact detection tasks across 17 databases containing diverse streaming artifacts in HD, UHD, HDR and HFR formats. This work represents an important step towards practical, actionable, and interpretable video quality assessment.

视频质量诊断评估专家模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。