arXiv:2608.05485cs.CV2026-08被引 2

VideoArgus用自动生成的评分标准统一评估视频生成与编辑,更贴近人类判断。

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing

论文配图:VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
图 1 · 摘自论文原文
  • 为每个输入生成专属评分标准,复用评估多个候选视频。
  • 在1260个视频上与人工评价相关性高于现有方法。
  • 适合需要可解释、一致评估的视频生成研究者使用。

视频生成的评估仍具挑战性,因现有基准依赖固定评价内容,仅覆盖部分生成与编辑场景,且评分缺乏证据支持。我们提出 VideoArgus,一个覆盖五种视频生成与编辑设置的统一评分框架。对每个输入实例,VideoArgus 生成一次输出无关、样本特定的评分标准,并复用于评估所有对应候选视频。该标准定义具体评分准则、打分规则、失败模式及证据计划,引导基于视觉语言模型(VLM)的问答与视觉工具生成有证据支持的评分、推理过程和诊断报告。我们进一步构建了 VideoArgus-Bench,包含 1,026 个精选输入实例,源自 653 张高质量图像和 416 段高质量视频,所有基准评分标准均已预生成、冻结并公开。在另一个人工对齐的 1,260 视频数据集上,VideoArgus 在所有五个任务中均展现出比对应基准专用评估器更高的输入内斯皮尔曼与肯德尔相关性。模型排名在不同评分标准生成方式和评估 VLM 骨干下也保持高度一致。全部代码与数据已公开。

原文摘要 · Abstract (English)

Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide limited evidence for their scores. We introduce VideoArgus, a unified rubric-grounded framework covering five video generation and editing settings. For each input instance, VideoArgus generates an output-blind, sample-specific rubric once and reuses it to evaluate all corresponding candidate videos. The rubric defines concrete criteria, scoring rules, failure modes, and evidence plans, which guide criterion-specific VLM QA and visual tools to produce evidence-grounded criterion scores, rationales, and a diagnostic report. We further construct VideoArgus-Bench, containing 1,026 curated input instances built from 653 high-quality images and 416 high-quality videos, with all benchmark rubrics pre-generated, frozen, and released. On a separate 1,260-video human-alignment set, VideoArgus achieves higher within-input Spearman and Kendall correlations with human judgments than the corresponding benchmark-specific evaluators across all five tasks. Model rankings also remain largely consistent across different rubric-generation and evaluation-VLM backbones. All code and data are released. Visit our project page: https://zzzmyyzeng.github.io/VideoArgus

视频评估评分标准生成质量VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。