提出可解释的视频生成指令遵循评估框架,自动诊断模型表现。
VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation

- 将复杂指令解析为时空有向无环图,实现依赖感知问答与短路诊断。
- 在223个长指令上测试14个模型,生成超3000段视频,结果可解释。
- 适合研究视频生成可控性、评估模型对复杂指令理解能力的人参考。
近期视频生成模型(VGMs)在视觉保真度上取得显著进展,但对长而复杂的指令遵循能力仍缺乏有效评估。现有评估方法多依赖简短且语义浅显的提示,约束原子性弱、时空依赖关系不足,且常需昂贵的人工评价或手工视觉流程,难以提供具体失败环节的诊断。为此,我们提出VGIF-Score——一个高度自动化且可解释的视频生成指令遵循评估框架。该框架包含两个互补模块:客观完成度分支将提示解析为时空有向无环图(ST-DAG),进行依赖感知问答与短路诊断;主观满意度分支则使用指令条件化的AutoRubric,评估电影感、视觉纯净度、运动流畅性及物理合理性。两者结合输出统一评分。我们在VGIF-Bench上实例化该框架,该基准包含223个结构复杂长指令,配以约4300个细粒度评估项。在超过3000段生成视频上对14个专有及开源VGM进行实验,结果表明VGIF-Score能可靠、可解释地评估视频生成的指令遵循能力。
原文摘要 · Abstract (English)
Recent video generation models (VGMs) have made substantial progress in visual fidelity, yet their ability to follow long, compositional instructions remains insufficiently evaluated. Existing evaluation protocols often rely on prompts that are short and semantically shallow, with limited atomic constraints and weak spatio-temporal dependencies. They also frequently depend on costly human evaluation or handcrafted vision pipelines, while providing little diagnostic insight into which instruction constraints succeed or fail. To address this gap, we propose VGIF-Score, a highly automated and interpretable framework for evaluating instruction following in video generation. VGIF-Score consists of two complementary components: an objective completion branch that parses prompts into a Spatio-Temporal Directed Acyclic Graph (ST-DAG) and performs dependency-aware QA with short-circuit diagnostics, and a subjective satisfaction branch that uses instruction-conditioned AutoRubric to assess cinematography, visual purity, motion smoothness, and physics adherence. Together, these components produce a unified score that captures both objective completion and perceptual satisfaction. We instantiate this framework on VGIF-Bench, a benchmark of 223 long, structurally entangled prompts paired with approximately 4.3K fine-grained evaluation items. Experiments on 14 proprietary and open-source VGMs across more than 3K generated videos show that VGIF-Score provides reliable, interpretable, and diagnostically useful evaluation of video generation instruction following. The code will be available at https://github.com/PRIS-CV/VGIF-SCORE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。