为电影级视频生成设计专业评估标准,真实还原导演创作逻辑。
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

- 从获奖影片中逆向提取真实拍摄脚本作为提示,确保每条指令对应实际镜头。
- 构建包含35项指标的电影语言评估体系,覆盖动态美学与多镜头连贯性。
- 自研专家级自动评测工具,可精准复现人类对模型的排序评价。
视频生成技术不断缩小与专业影视画面的差距,但现有基准仍使用网络或LLM生成的提示,并依赖未经训练的通用多模态模型评分。其评估体系仍停留在视觉质量、文本对齐和时序平滑等初级维度,未能反映电影制作中真正的专业电影语言标准。为此,我们推出FilmBench,一个基于电影学院传统电影语言体系的文本到视频(T2V)与参考到视频(R2V)基准,由北京电影学院及胡景数字媒体娱乐集团导演共同参与设计。该基准有三大核心:第一,提示源自100部获奖影片中的片段,经专业导演筛选,每条提示均对应真实实拍素材,且多数为多镜头(1,056/1,169),符合真实分镜脚本;第二,采用三级电影语言评估体系,涵盖3个维度、12个组件、35项(T2V)+3项(R2V独有)子指标;第三,开发专用专家级自动评估代理,并开源核心电影语言操作符套件(FilmOps)。在9个T2V与7个R2V模型上测试,评估结果与人类评分相关性达Spearman ρ = 0.95(T2V)与0.96(R2V),显著低于传统基准。结果显示动态美学表现明显不足,且弱模型在多镜头生成上性能断崖式下降。
原文摘要 · Abstract (English)
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。