arXiv:2608.09111cs.AI2026-08

用评分标准指导的AI模型,自动评估视频生成质量差异。

RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement

论文配图:RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement
图 1 · 摘自论文原文
  • 基于多模态大模型的评分准则,进行成对视频质量对比判断。
  • 覆盖150个文本到视频、100个图像到视频任务,收集超4500个生成视频。
  • 支持低成本新增模型评测,适合研究者与开发者快速比对性能。

AI视频生成技术快速发展并广泛商用,但先进模型间质量差异日益细微,传统视觉保真度和指令遵循等评估指标已难以区分。同时,人工评估需更高专业性与注意力,成本大幅上升。为此,我们提出RAVEN-Eval——一种基于评分标准引导的自动化评估框架,核心采用LMM-as-a-judge范式。通过自动任务筛选与质量过滤流程,RAVEN-Eval构建了150个文本到视频(T2V)任务与100个图像到视频(I2V)任务,系统收集超过4,500个AI生成视频。其核心机制为基于任务定制评分标准的LMM偏好判断,由多模态大模型执行成对比较。进一步引入锚点式模型插入方法,降低新模型评测成本。最终评估了20个高性能AIVGMs及13个LMM judges的判别能力,建立RAVEN-Eval Leaderboards。整体上,RAVEN-Eval为快速演进的AIVGM提供了可扩展、可信的自动化评估路径。

原文摘要 · Abstract (English)

AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following. Meanwhile, human evaluation now requires more expertise and sustained attention, substantially increasing annotation costs. This calls for automated evaluation that can reliably distinguish fine-grained differences among advanced AIVGMs with minimal human intervention. To address this challenge, we present RAVEN-Eval, a rubric-guided automated evaluation framework for AIVGMs, built primarily on the LMM-as-a-judge paradigm. Through an automatic task curation and quality-filtering pipeline, RAVEN-Eval curates 150 text-to-video~(T2V) tasks and 100 image-to-video~(I2V) tasks, and systematically collects more than 4,500 AIGVs. At its core, RAVEN-Eval adopts rubric-guided automated LMM preference judgement, in which LMM judges conduct pairwise comparisons according to task-specific rubrics. It further introduces an anchor-based model insertion approach to reduce the evaluation cost of incorporating new models. Finally, we evaluate 20 high-performance AIVGMs, as well as the judging capabilities of 13 LMM judges, and establish the RAVEN-Eval Leaderboards. Overall, RAVEN-Eval paves a scalable path for automatic and trustworthy evaluation of rapidly evolving AIVGMs.

视频生成自动评估多模态模型评分标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。