构建物理常识评测基准,检验视频生成模型是否懂基本物理规律。
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
- 设计160个覆盖27种物理定律的精细提示,评估视频生成对物理常识的理解。
- 通过多模型协同评估框架,自动化测试结果与人工反馈高度一致。
- 发现当前模型在动态场景下仍严重违背物理常识,需专门训练提升。
文本到视频(T2V)模型如Sora在可视化复杂提示方面取得显著进展,被视为构建通用世界模拟器的重要路径。认知心理学认为,实现该目标的基础在于理解直觉物理。然而,这些模型对直觉物理的准确表达能力尚未被充分探索。为填补这一空白,我们提出PhyGenBench——一个全面的物理生成评测基准,用于评估T2V生成中的物理常识正确性。PhyGenBench包含160个精心设计的提示,覆盖27种不同的物理定律,涵盖四个基础领域,可全面评估模型对物理常识的理解。同时,我们提出新型评估框架PhyGenEval,采用分层结构,结合先进的视觉-语言模型和大语言模型进行物理常识评估。通过PhyGenBench与PhyGenEval,可实现对T2V模型物理常识理解的大规模自动化评估,其结果与人类反馈高度一致。评估结果与深入分析表明,当前模型在生成符合物理常识的视频方面仍存在明显不足。仅通过扩大模型规模或使用提示工程无法完全解决PhyGenBench所揭示的挑战(如动态场景)。我们希望本研究能推动社区将物理常识学习置于这些模型发展的重要位置,超越娱乐应用。数据与代码将公开于https://github.com/OpenGVLab/PhyGenBench。
原文摘要 · Abstract (English)
Text-to-video (T2V) models like Sora have made significant strides in visualizing complex prompts, which is increasingly viewed as a promising path towards constructing the universal world simulator. Cognitive psychologists believe that the foundation for achieving this goal is the ability to understand intuitive physics. However, the capacity of these models to accurately represent intuitive physics remains largely unexplored. To bridge this gap, we introduce PhyGenBench, a comprehensive \textbf{Phy}sics \textbf{Gen}eration \textbf{Ben}chmark designed to evaluate physical commonsense correctness in T2V generation. PhyGenBench comprises 160 carefully crafted prompts across 27 distinct physical laws, spanning four fundamental domains, which could comprehensively assesses models' understanding of physical commonsense. Alongside PhyGenBench, we propose a novel evaluation framework called PhyGenEval. This framework employs a hierarchical evaluation structure utilizing appropriate advanced vision-language models and large language models to assess physical commonsense. Through PhyGenBench and PhyGenEval, we can conduct large-scale automated assessments of T2V models' understanding of physical commonsense, which align closely with human feedback. Our evaluation results and in-depth analysis demonstrate that current models struggle to generate videos that comply with physical commonsense. Moreover, simply scaling up models or employing prompt engineering techniques is insufficient to fully address the challenges presented by PhyGenBench (e.g., dynamic scenarios). We hope this study will inspire the community to prioritize the learning of physical commonsense in these models beyond entertainment applications. We will release the data and codes at https://github.com/OpenGVLab/PhyGenBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。