arXiv:2503.02341cs.CVcs.AI2025-03ICML被引 7

用多步推理评估文生视频,让机器评分更像人

GRADEO: Towards Human-Like Evaluation for Text-to-Video Generation via Multi-Step Reasoning

  • 基于人类标注构建多步推理评估数据集
  • 评分与人类判断更一致,准确率提升显著
  • 适合评估复杂场景下生成视频的质量

近期文本到视频生成模型取得显著进展,但有效评估仍具挑战。现有自动化评估指标缺乏高层语义理解与推理能力,难以解释且不贴近人类判断。为此,我们构建了GRADEO-Instruct数据集,包含来自10余种生成模型的3300个视频及由1.6万名人类标注转换的多步推理评估。进而提出GRADEO,首个专为视频评估设计的多步推理模型,可生成可解释的评分与分析。实验表明,该方法比现有方法更贴近人类评价。此外,基准测试揭示当前模型在符合人类推理和复杂现实场景方面仍有明显不足。

原文摘要 · Abstract (English)

Recent great advances in video generation models have demonstrated their potential to produce high-quality videos, bringing challenges to effective evaluation. Unlike human evaluation, existing automated evaluation metrics lack highlevel semantic understanding and reasoning capabilities for video, thus making them infeasible and unexplainable. To fill this gap, we curate GRADEO-Instruct, a multi-dimensional T2V evaluation instruction tuning dataset, including 3.3k videos from over 10 existing video generation models and multi-step reasoning assessments converted by 16k human annotations. We then introduce GRADEO, one of the first specifically designed video evaluation models, which grades AI-generated videos for explainable scores and assessments through multi-step reasoning. Experiments show that our method aligns better with human evaluations than existing methods. Furthermore, our benchmarking reveals that current video generation models struggle to produce content that aligns with human reasoning and complex real-world scenarios.

文生视频评估方法多步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。