arXiv:2503.23452cs.CV2025-03被引 26

用智能体系统更真实评估视频生成模型能力

VideoGen-Eval: Agent-based System for Video Generation Evaluation

  • 设计多智能体流程,结合语言与视觉模型动态评估视频内容
  • 在700个丰富提示下测试20多个模型,8个顶尖模型表现可靠
  • 评测结果与人工判断高度一致,适合研究者和开发者使用

视频生成技术的快速发展使得现有评估体系难以有效衡量先进模型的表现,主要问题在于提示过于简单、固定评估算子难以处理分布外(OOD)情况,以及指标与人类偏好不一致。为此,我们提出 VideoGen-Eval,一个基于智能体的评估系统,融合大语言模型的内容结构化、多模态大模型的内容判断及针对时序密集维度设计的补丁工具,实现动态、灵活且可扩展的视频生成评估。同时,我们构建了一个视频生成基准,包含700个结构化、内容丰富的提示(涵盖文本到视频与图像到视频),以及超过12,000个由20多个模型生成的视频,其中8个前沿模型被用于定量评估智能体与人类的一致性。大量实验验证了该智能体评估系统与人类偏好高度一致,能可靠完成评估,并展现出基准数据的多样性与丰富性。

原文摘要 · Abstract (English)

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators struggling with Out-of-Distribution (OOD) cases, and misalignment between computed metrics and human preferences. To bridge the gap, we propose VideoGen-Eval, an agent evaluation system that integrates LLM-based content structuring, MLLM-based content judgment, and patch tools designed for temporal-dense dimensions, to achieve a dynamic, flexible, and expandable video generation evaluation. Additionally, we introduce a video generation benchmark to evaluate existing cutting-edge models and verify the effectiveness of our evaluation system. It comprises 700 structured, content-rich prompts (both T2V and I2V) and over 12,000 videos generated by 20+ models, among them, 8 cutting-edge models are selected as quantitative evaluation for the agent and human. Extensive experiments validate that our proposed agent-based evaluation system demonstrates strong alignment with human preferences and reliably completes the evaluation, as well as the diversity and richness of the benchmark.

视频生成智能体评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。