arXiv:2505.22543cs.CVcs.AI2025-05被引 8

构建超大规模视频质量评估数据集,提升模型感知能力

Scaling-up Perceptual Video Quality Assessment

  • 设计高效框架构建人机协同的多模态指令数据集
  • 建成40万条指令的OmniVQA-Chat-400K数据集,领先现有规模
  • 适合视频质量分析、AI评测系统研发人员使用

数据扩展规律已被证实能显著提升大型多模态模型在各类下游任务中的性能。然而,在感知视频质量评估(VQA)领域,由于标注资源稀缺和数据集规模不足,扩展规律的潜力尚未被充分挖掘。为此,我们提出OmniVQA框架,高效构建高质量、人机协同的VQA多模态指令数据库(MIDB)。基于此,我们扩展生成了目前该领域最大的数据集OmniVQA-Chat-400K。研究聚焦技术与美学质量维度,提供丰富的上下文指令数据以支持细粒度评估知识。同时,我们构建了OmniVQA-MOS-20K数据集以增强模型量化评分能力。提出互补式训练策略,有效利用两类数据提升质量理解与评分能力。进一步提出OmniVQA-FG(fine-grain)基准测试,评估模型细粒度表现。结果表明,我们的模型在质量理解和评分任务中均达到当前最优水平。

原文摘要 · Abstract (English)

The data scaling law has been shown to significantly enhance the performance of large multi-modal models (LMMs) across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of scaling law remains unprecedented due to the scarcity of labeled resources and the insufficient scale of datasets. To address this, we propose \textbf{OmniVQA}, an efficient framework designed to efficiently build high-quality, human-in-the-loop VQA multi-modal instruction databases (MIDBs). We then scale up to create \textbf{OmniVQA-Chat-400K}, the largest MIDB in the VQA field concurrently. Our focus is on the technical and aesthetic quality dimensions, with abundant in-context instruction data to provide fine-grained VQA knowledge. Additionally, we have built the \textbf{OmniVQA-MOS-20K} dataset to enhance the model's quantitative quality rating capabilities. We then introduce a \textbf{complementary} training strategy that effectively leverages the knowledge from datasets for quality understanding and quality rating tasks. Furthermore, we propose the \textbf{OmniVQA-FG (fine-grain)-Benchmark} to evaluate the fine-grained performance of the models. Our results demonstrate that our models achieve state-of-the-art performance in both quality understanding and rating tasks.

视频质量评估多模态数据集构建大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。