arXiv:2507.13428cs.CVcs.AI2025-07被引 20

评测文生视频模型的物理合理性,发现多数模型仍难真实模拟物理规律。

"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

  • 构建多层级物理现象评测集,覆盖运动、能量守恒与复杂交互
  • 在1050个提示下测试12个模型,发现普遍存在物理错误
  • 引入反物理提示测试模型逻辑一致性,适合评估生成可靠性

视频生成模型在高质量、逼真内容生成上取得显著进展,但其对物理现象的准确模拟仍是关键且未解决的挑战。本文提出PhyWorldBench,一个全面的基准,用于评估文本到视频模型在遵循物理定律方面的表现。该基准涵盖从基础运动、能量守恒到刚体交互及人或动物运动等多种物理现象层次。此外,我们引入新的反物理类别,提示中故意违反现实物理规律,以评估模型能否在保持逻辑一致的前提下遵循此类指令。除大规模人工评估外,还设计了一种利用现有多模态大语言模型实现零样本物理真实性评估的简单有效方法。我们评估了12个最先进的文生视频模型,包括5个开源和5个专有模型,并进行详细对比分析。通过针对1050个精心设计的提示(涵盖基础、复合及反物理场景)的系统测试,识别出这些模型在遵循现实物理规律方面面临的关键挑战。进一步分析不同物理现象和提示类型下的性能表现,提出针对性提示设计建议以提升物理真实性。

原文摘要 · Abstract (English)

Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. This paper presents PhyWorldBench, a comprehensive benchmark designed to evaluate video generation models based on their adherence to the laws of physics. The benchmark covers multiple levels of physical phenomena, ranging from fundamental principles such as object motion and energy conservation to more complex scenarios involving rigid body interactions and human or animal motion. Additionally, we introduce a novel Anti-Physics category, where prompts intentionally violate real-world physics, enabling the assessment of whether models can follow such instructions while maintaining logical consistency. Besides large-scale human evaluation, we also design a simple yet effective method that utilizes current multimodal large language models to evaluate physics realism in a zero-shot fashion. We evaluate 12 state-of-the-art text-to-video generation models, including five open-source and five proprietary models, with detailed comparison and analysis. Through systematic testing across 1050 curated prompts spanning fundamental, composite, and anti-physics scenarios, we identify pivotal challenges these models face in adhering to real-world physics. We further examine their performance under diverse physical phenomena and prompt types, and derive targeted recommendations for crafting prompts that enhance fidelity to physical principles.

文生视频物理仿真基准评测生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。