arXiv:2412.04508eess.IVcs.CV2024-12综述被引 38

综述视频质量评估最新进展,涵盖数据集、算法与未来方向

Video Quality Assessment: A Comprehensive Survey

  • 系统梳理基于深度学习的视频质量评估方法
  • 分析大规模人类标注数据集对模型性能的关键作用
  • 适合从事多媒体质量评估研究的学者与工程师参考

视频质量评估(VQA)旨在预测视频质量,使其与人类主观感知高度一致。传统基于自然图像和/或视频统计的模型虽受视觉系统启发,但在真实用户生成内容(UGC)上的表现有限,尤其在近期从网络爬取的大规模多样化视频数据集中表现不佳。近年来,深度神经网络和大型多模态模型(LMMs)的进步显著提升了该任务性能,超越了以往手工设计的模型。众多基于深度学习的VQA模型相继提出,其发展得益于内容多样、规模庞大的人工标注数据库,这些数据库提供了可靠的感知质量基准数据。本文全面综述了近年来VQA算法的发展、基准测试研究及支持其发展的数据库,并分析了研究设计与算法架构中的开放性问题。项目代码见:https://github.com/taco-group/Video-Quality-Assessment-A-Comprehensive-Survey。

原文摘要 · Abstract (English)

Video quality assessment (VQA) is an important processing task, aiming at predicting the quality of videos in a manner highly consistent with human judgments of perceived quality. Traditional VQA models based on natural image and/or video statistics, which are inspired both by models of projected images of the real world and by dual models of the human visual system, deliver only limited prediction performances on real-world user-generated content (UGC), as exemplified in recent large-scale VQA databases containing large numbers of diverse video contents crawled from the web. Fortunately, recent advances in deep neural networks and Large Multimodality Models (LMMs) have enabled significant progress in solving this problem, yielding better results than prior handcrafted models. Numerous deep learning-based VQA models have been developed, with progress in this direction driven by the creation of content-diverse, large-scale human-labeled databases that supply ground truth psychometric video quality data. Here, we present a comprehensive survey of recent progress in the development of VQA algorithms and the benchmarking studies and databases that make them possible. We also analyze open research directions on study design and VQA algorithm architectures. Github link: https://github.com/taco-group/Video-Quality-Assessment-A-Comprehensive-Survey.

视频质量深度学习综述多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。