arXiv:2505.00337cs.LGcs.AI2025-05被引 35

首个基于物理定律的文生视频评测基准,揭示现有模型普遍违背基本物理规律。

T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

  • 构建12条物理法则测试集,结合人工评估与反事实实验检验生成视频真实性。
  • 所有模型平均合规率低于0.60,即使给出明确提示也难纠正物理错误。
  • 适合关注视频生成可信度、物理一致性研究的研究者和开发者。

近年来,文生视频生成模型在视觉质量和指令遵循方面取得显著进展,广泛应用于数字艺术创作与在线互动。然而,其对基本物理规律的遵守仍缺乏系统评估:许多输出违反刚体碰撞、能量守恒、重力动力学等约束,导致内容失真甚至误导。现有评测多依赖像素级自动指标,且仅针对简单生活场景,忽视人类判断与第一性原理物理。为此,我们提出 extbf{T2VPhysBench},一个基于第一性原理的评测基准,系统检验主流文生视频模型(开源与商业)是否遵守12项核心物理定律,涵盖牛顿力学、守恒律及现象学效应。该基准采用严谨的人工评估协议,并包含三项研究:(1) 综合合规性评估显示,所有模型在各定律类别中平均得分均低于0.60;(2) 提示消融实验表明,即使提供具体、针对性的物理提示,仍无法有效修复违规;(3) 反事实鲁棒性测试发现,当被明确指令违反物理规则时,模型仍会生成显式违背规则的视频。结果揭示当前架构存在持续性局限,并为未来实现真正物理感知的视频生成提供明确方向。

原文摘要 · Abstract (English)

Text-to-video generative models have made significant strides in recent years, producing high-quality videos that excel in both aesthetic appeal and accurate instruction following, and have become central to digital art creation and user engagement online. Yet, despite these advancements, their ability to respect fundamental physical laws remains largely untested: many outputs still violate basic constraints such as rigid-body collisions, energy conservation, and gravitational dynamics, resulting in unrealistic or even misleading content. Existing physical-evaluation benchmarks typically rely on automatic, pixel-level metrics applied to simplistic, life-scenario prompts, and thus overlook both human judgment and first-principles physics. To fill this gap, we introduce \textbf{T2VPhysBench}, a first-principled benchmark that systematically evaluates whether state-of-the-art text-to-video systems, both open-source and commercial, obey twelve core physical laws including Newtonian mechanics, conservation principles, and phenomenological effects. Our benchmark employs a rigorous human evaluation protocol and includes three targeted studies: (1) an overall compliance assessment showing that all models score below 0.60 on average in each law category; (2) a prompt-hint ablation revealing that even detailed, law-specific hints fail to remedy physics violations; and (3) a counterfactual robustness test demonstrating that models often generate videos that explicitly break physical rules when so instructed. The results expose persistent limitations in current architectures and offer concrete insights for guiding future research toward truly physics-aware video generation.

文生视频物理一致性评测基准生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。