arXiv:2505.23693cs.CVcs.AI2025-05ACL被引 7

评测大模型对AI生成视频的反馈能力,发现现有模型表现参差不齐。

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos

  • 构建新基准VF-Eval,涵盖四类AIGC视频评估任务。
  • 13个前沿多模态模型在该基准上表现不佳,最佳为GPT-4.1。
  • 可指导视频生成优化,提升模型与人类反馈的一致性。

多模态大模型(MLLMs)近年来被广泛研究用于自然视频问答,但现有评估大多聚焦于真实视频,忽视了如AI生成内容(AIGC)等合成视频。尽管部分视频生成工作使用MLLMs评估生成视频质量,但其对AIGC视频的解析能力仍缺乏系统探索。为此,我们提出新基准VF-Eval,包含四类任务:连贯性验证、错误感知、错误类型检测与推理评估,全面评估MLLMs在AIGC视频上的能力。我们在VF-Eval上评估了13个前沿MLLMs,发现即使表现最优的GPT-4.1也未能在所有任务中保持一致高分,凸显该基准的挑战性。此外,为探究其实际应用价值,我们开展实验RePrompt,证明更贴近人类反馈的对齐可有效提升视频生成效果。

原文摘要 · Abstract (English)

MLLMs have been widely studied for video question answering recently. However, most existing assessments focus on natural videos, overlooking synthetic videos, such as AI-generated content (AIGC). Meanwhile, some works in video generation rely on MLLMs to evaluate the quality of generated videos, but the capabilities of MLLMs on interpreting AIGC videos remain largely underexplored. To address this, we propose a new benchmark, VF-Eval, which introduces four tasks-coherence validation, error awareness, error type detection, and reasoning evaluation-to comprehensively evaluate the abilities of MLLMs on AIGC videos. We evaluate 13 frontier MLLMs on VF-Eval and find that even the best-performing model, GPT-4.1, struggles to achieve consistently good performance across all tasks. This highlights the challenging nature of our benchmark. Additionally, to investigate the practical applications of VF-Eval in improving video generation, we conduct an experiment, RePrompt, demonstrating that aligning MLLMs more closely with human feedback can benefit video generation.

多模态AIGC评测基准视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。