构建多维视频生成评估基准,用大模型打分更准更省成本。
WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models

- 用Likert量表+原分辨率帧,替代二元问答,提升评估精度。
- 16维联合测试,仅需50个提示就覆盖所有维度,大幅降低生成成本。
- 大模型评分与人类偏好高度一致,可替代人工评测,适合研究者使用。
生成式视频模型的评估仍是开放问题。传统指标如SSIM和PSNR过度关注像素保真度,而FVD偏向分布纹理而非物理合理性。基于二元视觉问答(VQA)的基准如VBench 2.0易受‘是’偏见影响,且依赖低分辨率审核器,难以捕捉时间性错误。此外,其提示仅聚焦单一维度,导致所需视频数量激增,仍无法保证结果可靠性。WorldJen直接解决这些问题:将二元VQA替换为由视觉语言模型(VLM)在原生分辨率下评分的李克特量表问卷;通过对抗性策划的提示,可同时激发最多16个质量维度。框架包含两项核心贡献:首先,开展盲测人类偏好实验,收集7名标注者对50个策划提示×6个先进视频模型的2,696组两两对比注释,实现66.9%平均标注者一致性,并建立具有三层结构的人类真实布拉德利-泰瑞(BT)评分;其次,使用特定提示、特定维度的李克特问卷(每维10题,共47,160条评分),由VLM作为评判引擎独立复现该三层结构。该方法在斯皮尔曼等级相关系数上达到ρ̂=1.000,p=0.0014,表明与人类结果完全一致。六项消融实验验证了评估框架的鲁棒性。
原文摘要 · Abstract (English)
Evaluating generative video models remains an open problem. Reference-based metrics such as Structural Similarity Index Measure (SSIM) and Peak Signal to Noise Ratio (PSNR) reward pixel fidelity over semantic correctness, while Frechet Video Distance (FVD) favors distributional textures over physical plausibility. Binary Visual Question Answering (VQA) based benchmarks like VBench~2.0 are prone to yes-bias and rely on low-resolution auditors that miss temporal failures. Moreover, their prompts target a single dimension at a time, multiplying the number of videos required while still not guaranteeing reliable results. WorldJen addresses these limitations directly. Binary VQA is replaced with Likert-scale questionnaires graded by a VLM that receives frames at native video resolution. Video generation costs are addressed by using adversarially curated prompts that are designed to exercise up to 16 quality dimensions simultaneously. The framework is built around two interlocking contributions. First, A blind human preference study is conducted, accumulating (2,696 pairwise annotations from 7 annotators with 100% pair coverage over 50 of the curated prompts $\times$ 6 state-of-the-art video models. A mean inter-annotator agreement of 66.9% is achieved and the study establishes a human ground-truth Bradley-Terry (BT) rating with a three-tier structure. Second, A VLM-as-a-judge evaluation engine using prompt-specific, dimension-specific Likert questionnaires (10 questions per dimension, 47,160 scored responses) judges the videos and reproduces the human-established three-tier BT rating structure independently. The VLM achieves a Spearman $\hatρ=1.000,~p=0.0014$ that is interpreted as tier agreement with the human results. Six focused ablation studies validate the robustness of the VLM evaluation framework. Project page: https://moonmath.ai/worldjen/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。