arXiv:2503.21755cs.CV2025-03被引 257

VBench-2.0新基准评估视频生成的内在真实感,超越表面流畅性。

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

论文配图:VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
图 1 · 摘自论文原文
  • 构建五维评估体系,涵盖人体、物理、常识等内在真实性维度。
  • 融合通用模型与专用检测方法,实现自动化、高精度评估。
  • 适合关注视频生成真实性和可信度的研究者与开发者。

视频生成已从早期不切实际的输出发展为视觉逼真、时间连贯的成果。现有基准如VBench主要评估表面真实感,包括帧质量、时间一致性和提示遵循度。然而这些仍局限于表层表现,未能衡量生成内容是否符合真实世界规律。尽管当前模型在这些指标上表现优异,但依然难以生成真正符合物理法则、常识推理、解剖结构和组合完整性的视频。为推动视频生成迈向“世界模型”级别,必须突破表面真实感,实现内在真实感。为此,我们提出VBench-2.0——下一代视频生成评估基准,系统评估五大核心维度:人体真实感、可控性、创造力、物理合理性与常识性,每项进一步细分为多个子能力。评估框架结合顶尖视觉语言模型(VLM)、大语言模型(LLM)及专用于视频生成的异常检测方法,并通过大规模人工标注确保与人类判断对齐。该基准旨在引领下一代视频生成模型向内在真实感迈进。

原文摘要 · Abstract (English)

Video generation has advanced significantly, evolving from producing unrealistic outputs to generating videos that appear visually convincing and temporally coherent. To evaluate these video generative models, benchmarks such as VBench have been developed to assess their faithfulness, measuring factors like per-frame aesthetics, temporal consistency, and basic prompt adherence. However, these aspects mainly represent superficial faithfulness, which focus on whether the video appears visually convincing rather than whether it adheres to real-world principles. While recent models perform increasingly well on these metrics, they still struggle to generate videos that are not just visually plausible but fundamentally realistic. To achieve real "world models" through video generation, the next frontier lies in intrinsic faithfulness to ensure that generated videos adhere to physical laws, commonsense reasoning, anatomical correctness, and compositional integrity. Achieving this level of realism is essential for applications such as AI-assisted filmmaking and simulated world modeling. To bridge this gap, we introduce VBench-2.0, a next-generation benchmark designed to automatically evaluate video generative models for their intrinsic faithfulness. VBench-2.0 assesses five key dimensions: Human Fidelity, Controllability, Creativity, Physics, and Commonsense, each further broken down into fine-grained capabilities. Tailored to individual dimensions, our evaluation framework integrates generalists such as SOTA VLMs and LLMs, and specialists, including anomaly detection methods proposed for video generation. We conduct extensive human annotations to ensure evaluation alignment with human judgment. By pushing beyond superficial faithfulness toward intrinsic faithfulness, VBench-2.0 aims to set a new standard for the next generation of video generative models in pursuit of intrinsic faithfulness.

视频生成真实感评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。