测试视频生成模型对物理常识的理解能力,发现它们常犯常识性错误。
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
- 设计新基准测试,聚焦工具使用和材料属性等物理常识
- 用383个提示生成视频,评估结果在10%以上存在物理错误
- 适合关注视频生成可信度的研究者与开发者
文本到视频(T2V)生成技术虽能合成视觉逼真、时间连贯的视频,但往往缺乏基本物理常识,导致输出违背因果关系、物体行为和工具使用等直观预期。为此,我们提出PhysVidBench,一个用于评估T2V系统物理推理能力的基准。该基准包含383个精心设计的提示,强调工具使用、材料属性和过程性交互,这些领域对物理合理性至关重要。针对每个提示,我们使用多种前沿模型生成视频,并采用三阶段评估流程:(1) 从提示中提炼出基于物理的问题;(2) 用视觉-语言模型对生成视频进行描述;(3) 让语言模型仅根据描述回答多个涉及物理的问题。此间接策略有效规避了直接视频评估中的常见幻觉问题。通过突出工具可用性和工具中介动作,PhysVidBench为评估生成视频模型的物理常识提供了结构化、可解释的框架。
原文摘要 · Abstract (English)
Recent progress in text-to-video (T2V) generation has enabled the synthesis of visually compelling and temporally coherent videos from natural language. However, these models often fall short in basic physical commonsense, producing outputs that violate intuitive expectations around causality, object behavior, and tool use. Addressing this gap, we present PhysVidBench, a benchmark designed to evaluate the physical reasoning capabilities of T2V systems. The benchmark includes 383 carefully curated prompts, emphasizing tool use, material properties, and procedural interactions, and domains where physical plausibility is crucial. For each prompt, we generate videos using diverse state-of-the-art models and adopt a three-stage evaluation pipeline: (1) formulate grounded physics questions from the prompt, (2) caption the generated video with a vision-language model, and (3) task a language model to answer several physics-involved questions using only the caption. This indirect strategy circumvents common hallucination issues in direct video-based evaluation. By highlighting affordances and tool-mediated actions, areas overlooked in current T2V evaluations, PhysVidBench provides a structured, interpretable framework for assessing physical commonsense in generative video models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。