arXiv:2503.06800cs.CV2025-03被引 126

评测视频生成模型对物理常识的遵守程度,发现顶尖模型仅22%达标。

VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

  • 构建200个动作的精细视频生成评估集,聚焦物理规则
  • 人类评估显示最优模型在困难子集上联合表现仅22%
  • 揭示模型在质量与动量守恒等物理规律上严重不足

大规模视频生成模型虽能生成多样真实视频,但其在现实动作中对物理常识的遵循仍不明确(如打网球、后空翻)。现有基准存在规模小、缺乏人工评估、仿真到现实差距大及缺乏细粒度物理规则分析等问题。为此,我们提出VideoPhy-2,一个以动作为中心的生成视频物理常识评估数据集。我们收集了200个多样化动作及详细提示,用于现代生成模型的视频合成。通过人工评估,我们考察生成视频的语义一致性、物理常识性及物理规则的合理嵌入。结果显示存在显著缺陷:即使最佳模型在VideoPhy-2的困难子集上也仅达22%的联合性能(即高语义与物理常识一致)。模型尤其难以处理质量与动量守恒等守恒定律。此外,我们还训练了VideoPhy-AutoEval,一种自动评估器,可快速可靠地对本数据集进行评估。总体而言,VideoPhy-2提供了一个严格基准,暴露了视频生成模型在物理合理性方面的关键差距,并为未来研究指明方向。数据与代码见https://videophy2.github.io/。

原文摘要 · Abstract (English)

Large-scale video generative models, capable of creating realistic videos of diverse visual concepts, are strong candidates for general-purpose physical world simulators. However, their adherence to physical commonsense across real-world actions remains unclear (e.g., playing tennis, backflip). Existing benchmarks suffer from limitations such as limited size, lack of human evaluation, sim-to-real gaps, and absence of fine-grained physical rule analysis. To address this, we introduce VideoPhy-2, an action-centric dataset for evaluating physical commonsense in generated videos. We curate 200 diverse actions and detailed prompts for video synthesis from modern generative models. We perform human evaluation that assesses semantic adherence, physical commonsense, and grounding of physical rules in the generated videos. Our findings reveal major shortcomings, with even the best model achieving only 22% joint performance (i.e., high semantic and physical commonsense adherence) on the hard subset of VideoPhy-2. We find that the models particularly struggle with conservation laws like mass and momentum. Finally, we also train VideoPhy-AutoEval, an automatic evaluator for fast, reliable assessment on our dataset. Overall, VideoPhy-2 serves as a rigorous benchmark, exposing critical gaps in video generative models and guiding future research in physically-grounded video generation. The data and code is available at https://videophy2.github.io/.

视频生成物理常识评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。