arXiv:2603.29186cs.CVcs.AI2026-03中稿 · CVPR被引 2

构建合成长视频评估基准,检验文本生成视频的评价系统可靠性

SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generation

  • 基于成对比较框架,合成10类可控退化视频对
  • 人类可准确识别优劣视频,正确率达84.7%-96.8%
  • 9类场景下现有评估系统表现低于人类,暴露评测短板

本文提出合成长视频元评估基准(SLVMEval),用于评估文本到视频(T2V)生成系统的评价能力。该基准针对长达10,486秒(约3小时)的视频,关注系统在人类易判断质量的场景中是否具备准确评估能力。采用基于成对比较的元评估框架,利用密集视频标注数据集,合成多类可控退化的“高质量与低质量”视频对,覆盖10个不同方面。通过众包筛选,仅保留人类可清晰感知差异的视频对,构建有效测试集。利用该测试集评估现有评估系统的排序可靠性。实验结果表明,人类在识别更优长视频时准确率达84.7%-96.8%;在10个方面中的9个,现有系统准确率低于人类水平,揭示当前文本生成长视频评价体系的显著缺陷。

原文摘要 · Abstract (English)

This paper proposes the synthetic long-video meta-evaluation (SLVMEval), a benchmark for meta-evaluating text-to-video (T2V) evaluation systems. The proposed SLVMEval benchmark focuses on assessing these systems on videos of up to 10,486 s (approximately 3 h). The benchmark targets a fundamental requirement, namely, whether the systems can accurately assess video quality in settings that are easy for humans to assess. We adopt a pairwise comparison-based meta-evaluation framework. Building on dense video-captioning datasets, we synthetically degrade source videos to create controlled "high-quality versus low-quality" pairs across 10 distinct aspects. Then, we employ crowdsourcing to filter and retain only those pairs in which the degradation is clearly perceptible, thereby establishing an effective final testbed. Using this testbed, we assess the reliability of existing evaluation systems in ranking these pairs. Experimental results demonstrate that human evaluators can identify the better long video with 84.7%-96.8% accuracy, and in nine of the 10 aspects, the accuracy of these systems falls short of human assessment, revealing weaknesses in text-to-long-video evaluation.

视频生成评估基准元评估长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。