arXiv:2507.01255cs.CV2025-07被引 2

AI视频评估新模型,能打分还能生成多维度评语。

AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation

  • 统一模型同时输出评分与多方面语言评语。
  • 在2500个AI生成视频上实现最高人类评价一致性。
  • 适合需要可解释评估的视频生成研究与应用者。

AI生成视频的快速发展催生了对鲁棒且可解释评估框架的迫切需求。现有度量标准仅提供数值评分,缺乏解释性评论,导致可解释性差且与人工评价不一致。为此,我们提出AIGVE-MACS,一种统一的AI生成视频评估(AIGVE)模型,不仅能输出数值评分,还能提供多方面语言评论反馈。核心是AIGVE-BENCH 2,一个包含2,500个AI生成视频及22,500条人工标注的详细评论和数值评分的大规模基准,覆盖九个关键评估维度。基于该基准,AIGVE-MACS融合最新视觉-语言模型,采用新型逐标记加权损失和动态帧采样策略,显著提升与人工评价的一致性。在监督与零样本基准上的综合实验表明,AIGVE-MACS在评分相关性和评论质量上均达到领先水平,显著优于包括GPT-4o和VideoScore在内的基线方法。此外,我们进一步展示了一个多智能体优化框架,利用AIGVE-MACS的反馈驱动视频生成的迭代改进,实现53.5%的质量提升。该工作建立了全面、与人类对齐的AI生成视频评估新范式。我们已在Hugging Face公开发布AIGVE-BENCH 2和AIGVE-MACS。

原文摘要 · Abstract (English)

The rapid advancement of AI-generated video models has created a pressing need for robust and interpretable evaluation frameworks. Existing metrics are limited to producing numerical scores without explanatory comments, resulting in low interpretability and human evaluation alignment. To address those challenges, we introduce AIGVE-MACS, a unified model for AI-Generated Video Evaluation(AIGVE), which can provide not only numerical scores but also multi-aspect language comment feedback in evaluating these generated videos. Central to our approach is AIGVE-BENCH 2, a large-scale benchmark comprising 2,500 AI-generated videos and 22,500 human-annotated detailed comments and numerical scores across nine critical evaluation aspects. Leveraging AIGVE-BENCH 2, AIGVE-MACS incorporates recent Vision-Language Models with a novel token-wise weighted loss and a dynamic frame sampling strategy to better align with human evaluators. Comprehensive experiments across supervised and zero-shot benchmarks demonstrate that AIGVE-MACS achieves state-of-the-art performance in both scoring correlation and comment quality, significantly outperforming prior baselines including GPT-4o and VideoScore. In addition, we further showcase a multi-agent refinement framework where feedback from AIGVE-MACS drives iterative improvements in video generation, leading to 53.5% quality enhancement. This work establishes a new paradigm for comprehensive, human-aligned evaluation of AI-generated videos. We release the AIGVE-BENCH 2 and AIGVE-MACS at https://huggingface.co/xiaoliux/AIGVE-MACS.

视频评估多模态AI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。