arXiv:2511.04628cs.CV2025-11

用视频时序信息做无参考质量评估,无需人工评分标签。

NovisVQ: A Streaming Convolutional Neural Network for No-Reference Opinion-Unaware Frame Quality Assessment

  • 基于时序卷积网络,从受损视频直接预测质量分数。
  • 在多种失真下表现优于图像基线,与全参考指标相关性更高。
  • 适合实时视频系统,无需参考视频和人类评分数据。

视频质量评估对计算机视觉任务至关重要,但现有方法存在明显局限:全参考(FR)指标需干净参考视频,多数无参考(NR)模型依赖昂贵的人工评分训练。此外,多数无参考评估方法为图像级,忽略视频中对目标检测至关重要的时序上下文。本文提出一种可扩展的流式无参考、无主观意见的视频质量评估模型。该模型利用DAVIS数据集的合成退化视频,训练一个具备时序感知能力的卷积架构,直接从受损视频预测全参考指标(LPIPS、PSNR、SSIM),推理时不需参考视频。实验表明,该流式方法相比自建图像基线,在多样退化场景下泛化能力更强,凸显时序建模对可扩展视频质量评估的价值。同时,其与全参考指标的相关性高于广泛使用的BRISQUE(基于主观评价的图像评估基线),验证了时序、无主观方法的有效性。

原文摘要 · Abstract (English)

Video quality assessment (VQA) is vital for computer vision tasks, but existing approaches face major limitations: full-reference (FR) metrics require clean reference videos, and most no-reference (NR) models depend on training on costly human opinion labels. Moreover, most opinion-unaware NR methods are image-based, ignoring temporal context critical for video object detection. In this work, we present a scalable, streaming-based VQA model that is both no-reference and opinion-unaware. Our model leverages synthetic degradations of the DAVIS dataset, training a temporal-aware convolutional architecture to predict FR metrics (LPIPS , PSNR, SSIM) directly from degraded video, without references at inference. We show that our streaming approach outperforms our own image-based baseline by generalizing across diverse degradations, underscoring the value of temporal modeling for scalable VQA in real-world vision systems. Additionally, we demonstrate that our model achieves higher correlation with full-reference metrics compared to BRISQUE, a widely-used opinion-aware image quality assessment baseline, validating the effectiveness of our temporal, opinion-unaware approach.

视频质量无参考时序建模流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。