用视觉问答框架提升视频质量评估,实现更全面的感知理解。
VQA$^2$: Visual Question Answering for Video Quality Assessment
- 构建首个面向视频质量的视觉问答数据集,含15万+问答对。
- 提出VQA2系列模型,通过时空特征融合显著提升质量感知能力。
- 模型在质量理解任务上超越GPT-4o,适合多媒体质量分析场景。
大型多模态模型(LMMs)的兴起为计算机视觉带来了新范式,将各类任务统一为视觉问答框架。视频质量评估(VQA)作为低层视觉感知的经典领域,最初聚焦于量化评分。然而,随着LMMs的发展,正向更全面的视觉质量理解演进。已有研究在图像领域证明,视觉问答(VQA)可显著提升低层视觉质量评价效果。但视频领域的相关工作仍为空白,存在巨大改进空间。为此,我们提出了VQA²指令数据集——首个专注于视频质量评估的视觉问答指令数据集,包含3个子集和多种视频类型,共包含157,755条指令问答对。在此基础上,我们提出了VQA²系列模型,通过交错处理视觉与运动令牌,增强对视频时空质量细节的感知。我们在视频质量评分与理解任务上进行了大量实验,结果表明,VQA²系列模型在两项任务中均表现优异。特别地,最终模型VQA²-Assistant在视觉质量理解任务上超越了知名的GPT-4o,同时在评分任务中保持强竞争力。本工作为低层视频质量评估与理解与LMMs的融合提供了基础与可行路径。
原文摘要 · Abstract (English)
The advent and proliferation of large multi-modal models (LMMs) have introduced new paradigms to computer vision, transforming various tasks into a unified visual question answering framework. Video Quality Assessment (VQA), a classic field in low-level visual perception, focused initially on quantitative video quality scoring. However, driven by advances in LMMs, it is now progressing toward more holistic visual quality understanding tasks. Recent studies in the image domain have demonstrated that Visual Question Answering (VQA) can markedly enhance low-level visual quality evaluation. Nevertheless, related work has not been explored in the video domain, leaving substantial room for improvement. To address this gap, we introduce the VQA2 Instruction Dataset - the first visual question answering instruction dataset that focuses on video quality assessment. This dataset consists of 3 subsets and covers various video types, containing 157,755 instruction question-answer pairs. Then, leveraging this foundation, we present the VQA2 series models. The VQA2 series models interleave visual and motion tokens to enhance the perception of spatial-temporal quality details in videos. We conduct extensive experiments on video quality scoring and understanding tasks, and results demonstrate that the VQA2series models achieve excellent performance in both tasks. Notably, our final model, the VQA2-Assistant, exceeds the renowned GPT-4o in visual quality understanding tasks while maintaining strong competitiveness in quality scoring tasks. Our work provides a foundation and feasible approach for integrating low-level video quality assessment and understanding with LMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。