arXiv:2503.09949cs.CV2025-03NeurIPS被引 7

用多模态大模型统一评估AI生成视频,效果超越专用方法。

UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?

  • 用多模态大模型直接评估视频,无需为每项指标训练专用模型。
  • 在15个维度上对比,顶尖模型表现接近人类,显著优于现有方法。
  • 提出UVE-Bench基准,支持细粒度、全面的视频质量评估,适合评测者参考。

随着视频生成模型(VGMs)的快速发展,构建可靠且全面的自动评估指标对AI生成视频(AIGVs)至关重要。现有方法要么使用为其他任务优化的现成模型,要么依赖人工标注数据训练专用评估器,这些方法受限于特定评估维度,难以适应日益增长的细粒度与综合性评估需求。本文探讨将多模态大语言模型(MLLMs)作为统一评估器的可行性,利用其强大的视觉感知与语言理解能力。为此,我们构建了名为UVE-Bench的基准,收集了来自先进VGM的视频,并提供跨15个评估维度的成对人类偏好标注。基于UVE-Bench,我们系统评估了18个MLLMs。实验证明,尽管先进MLLMs(如Qwen2VL-72B和InternVL2.5-78B)仍不及人类评估者,但在统一评估中展现出良好潜力,显著优于现有专用评估方法。此外,我们深入分析了影响MLLM评估性能的关键设计因素,为未来AIGV评估研究提供了重要启示。

原文摘要 · Abstract (English)

With the rapid growth of video generative models (VGMs), it is essential to develop reliable and comprehensive automatic metrics for AI-generated videos (AIGVs). Existing methods either use off-the-shelf models optimized for other tasks or rely on human assessment data to train specialized evaluators. These approaches are constrained to specific evaluation aspects and are difficult to scale with the increasing demands for finer-grained and more comprehensive evaluations. To address this issue, this work investigates the feasibility of using multimodal large language models (MLLMs) as a unified evaluator for AIGVs, leveraging their strong visual perception and language understanding capabilities. To evaluate the performance of automatic metrics in unified AIGV evaluation, we introduce a benchmark called UVE-Bench. UVE-Bench collects videos generated by state-of-the-art VGMs and provides pairwise human preference annotations across 15 evaluation aspects. Using UVE-Bench, we extensively evaluate 18 MLLMs. Our empirical results suggest that while advanced MLLMs (e.g., Qwen2VL-72B and InternVL2.5-78B) still lag behind human evaluators, they demonstrate promising ability in unified AIGV evaluation, significantly surpassing existing specialized evaluation methods. Additionally, we conduct an in-depth analysis of key design choices that impact the performance of MLLM-driven evaluators, offering valuable insights for future research on AIGV evaluation.

视频评估多模态模型自动评测AI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。