arXiv:2507.18107cs.CV2025-07被引 21

首个评估文本生成视频世界知识能力的基准,揭示主流模型普遍缺乏常识理解。

T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation

  • 构建涵盖6大类、60子类的1200个提示的系统性评测框架
  • 发现10个主流T2V模型在常识一致性上表现不佳,多数生成内容错误
  • 结合人工评分与视觉语言模型自动评估,兼顾主观偏好与可扩展性

文本到视频(T2V)模型在生成视觉合理场景方面表现突出,但其利用世界知识确保语义一致性和事实准确性的能力仍缺乏系统研究。为此,我们提出T2VWorldBench,首个针对T2V模型世界知识生成能力的系统性评估框架,覆盖物理、自然、行为、文化、因果关系和物体等6大类别,包含60个子类和1200个提示,涵盖广泛领域。为兼顾人类偏好与可扩展性评估,该基准融合人工评价与基于视觉语言模型(VLMs)的自动化评估。我们评估了当前最先进的10个T2V模型,涵盖开源与商用模型,发现大多数模型无法理解世界知识,生成的视频存在事实性错误。这一结果揭示了现有T2V模型在利用世界知识方面的关键短板,为构建具备强常识推理与真实生成能力的模型提供了重要研究方向与切入点。

原文摘要 · Abstract (English)

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic consistency and factual accuracy remains largely understudied. In response to this challenge, we propose T2VWorldBench, the first systematic evaluation framework for evaluating the world knowledge generation abilities of text-to-video models, covering 6 major categories, 60 subcategories, and 1,200 prompts across a wide range of domains, including physics, nature, activity, culture, causality, and object. To address both human preference and scalable evaluation, our benchmark incorporates both human evaluation and automated evaluation using vision-language models (VLMs). We evaluated the 10 most advanced text-to-video models currently available, ranging from open source to commercial models, and found that most models are unable to understand world knowledge and generate truly correct videos. These findings point out a critical gap in the capability of current text-to-video models to leverage world knowledge, providing valuable research opportunities and entry points for constructing models with robust capabilities for commonsense reasoning and factual generation.

文本生成视频世界知识评估基准常识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。