arXiv:2509.11589cs.CV2025-09被引 4

构建6.8万视频多维度质量数据集,提升视频评估可解释性。

MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment

  • 基于7个维度标注超6.8万视频,附带推理链条增强可解释性。
  • 在多个公开基准上实现当前最佳性能,零样本泛化能力显著提升。
  • 适合视频生成、评估模型研发者,推动高质量视频筛选研究。

随着Sora等视频生成模型的快速发展,从大规模预训练数据集中筛选高质量视频的视频质量评估(VQA)变得日益关键。传统VQA方法通常仅输出单一数值评分,缺乏全面性和可解释性。为此,我们提出MVQA-68K,一个包含超过68,000条精心标注视频的多维度VQA数据集,覆盖整体美学、镜头运动、动态程度、纹理细节、构图、视觉质量和事实一致性共七个核心质量维度。每条标注均包含详细的链式思维推理过程,以促进可解释性与全面理解。大量实验表明,使用MVQA-68K可显著提升多种多模态大语言模型(MLLMs)在VQA任务上的表现,在内部测试集(图1)及公开基准如LSVQ-test、LSVQ-1080p和LIVE-VQC上均达到领先水平。同时,在VQA训练中引入显式推理过程,能大幅提升零样本泛化能力。代码与数据集将开源至GitHub:https://github.com/Controller01-ai/MVQA-68K。

原文摘要 · Abstract (English)

With the rapid advancement of video generation models such as Sora, video quality assessment (VQA) is becoming increasingly crucial for selecting high-quality videos from large-scale datasets used in pre-training. Traditional VQA methods, typically producing single numerical scores, often lack comprehensiveness and interpretability. To address these challenges, we introduce MVQA-68K, a novel multi-dimensional VQA dataset comprising over 68,000 carefully annotated videos, covering seven essential quality dimensions: overall aesthetics, camera movement, dynamic degree, texture detail, composition, visual quality, and factual consistency. Each annotation includes detailed chain-of-thought reasoning to facilitate interpretability and comprehensive understanding. Extensive experiments demonstrate that MVQA-68K significantly enhances the performance of various multimodal large language models (MLLMs) on the VQA task, achieving state-of-the-art results not only on our internal test set (Fig.1) but also on public benchmarks including LSVQ-test, LSVQ-1080p, and LIVE-VQC. Meantime, incorporating explicit reasoning process during VQA training substantially boosts the zero-shot generalization. Code and dataset will be available at github: https://github.com/Controller01-ai/MVQA-68K

视频评估多维标注可解释性MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。