通过分析视频帧间差异,实现无参考视频质量评估的高效高精度检测。
DIVA-VQA: Detecting Inter-frame Variations in UGC Video Quality
- 基于帧间变化驱动时空碎片化,分层级分析质量敏感区域。
- 在五个UGC数据集上平均相关性达0.898(DIVA-VQA-L),位居前列。
- 轻量级设计,运行效率优于多数现有方法,适合大规模视频监控。
用户生成内容(UGC)的快速增长推动了无参考(NR)感知视频质量评估(VQA)的研究需求。在社交媒体和流媒体应用中,由于缺乏原始参考视频,NR-VQA成为大规模视频质量监控的关键。本文提出一种基于帧间差异驱动的时空碎片化新型NR-VQA模型。该模型利用帧间差异,逐层分析帧、块及碎片化帧的质量敏感区域,融合对齐残差的帧、碎片残差与碎片化帧,有效捕捉全局与局部信息。模型同时提取2D与3D特征以表征时空变化。在五个UGC数据集上的实验表明,所提方法平均排名位列前二:DIVA-VQA-L相关性为0.898,DIVA-VQA-B为0.886。性能提升的同时保持低运行复杂度,其中DIVA-VQA-B在速度上排名第一,DIVA-VQA-L排名第三,优于当前最快的方法。代码与模型已公开于:https://github.com/xinyiW915/DIVA-VQA。
原文摘要 · Abstract (English)
The rapid growth of user-generated (video) content (UGC) has driven increased demand for research on no-reference (NR) perceptual video quality assessment (VQA). NR-VQA is a key component for large-scale video quality monitoring in social media and streaming applications where a pristine reference is not available. This paper proposes a novel NR-VQA model based on spatio-temporal fragmentation driven by inter-frame variations. By leveraging these inter-frame differences, the model progressively analyses quality-sensitive regions at multiple levels: frames, patches, and fragmented frames. It integrates frames, fragmented residuals, and fragmented frames aligned with residuals to effectively capture global and local information. The model extracts both 2D and 3D features in order to characterize these spatio-temporal variations. Experiments conducted on five UGC datasets and against state-of-the-art models ranked our proposed method among the top 2 in terms of average rank correlation (DIVA-VQA-L: 0.898 and DIVA-VQA-B: 0.886). The improved performance is offered at a low runtime complexity, with DIVA-VQA-B ranked top and DIVA-VQA-L third on average compared to the fastest existing NR-VQA method. Code and models are publicly available at: https://github.com/xinyiW915/DIVA-VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。