用生成质量评估视频模型的物理理解能力,更准更省事。
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
- 不训练模型,直接用去噪目标判断视频是否符合物理规律。
- 在12种场景下,评估结果与人类判断高度一致,优于现有方法。
- 发现模型越大、推理越精细,物理理解越强,但复杂动态仍难搞定。
视频扩散模型中的直观物理理解对构建通用且符合物理的世界模拟器至关重要,但准确评估其能力仍具挑战,因生成结果的物理正确性与视觉外观难以分离。为此,我们提出LikePhys,一种无需训练的方法:通过一个精心构建的有效-无效视频对数据集,利用去噪目标作为基于ELBO的似然近似,区分物理上合理与不可能的视频。我们在涵盖四大物理领域的十二种场景构成的基准上进行测试,结果显示,所提出的可置信度偏好误差(PPE)指标与人类偏好高度一致,优于现有最先进评估基线。我们系统性地评估了当前视频扩散模型的直观物理理解能力。研究进一步分析了模型设计与推断设置对物理理解的影响,并揭示了不同物理规律下的领域特异性能力差异。实验证明,尽管现有模型在复杂和混沌动力学中表现不佳,但随着模型容量和推断设置的提升,物理理解能力呈现明显改善趋势。
原文摘要 · Abstract (English)
Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in generation. To the end, we introduce LikePhys, a training-free method that evaluates intuitive physics in video diffusion models by distinguishing physically valid and impossible videos using the denoising objective as an ELBO-based likelihood surrogate on a curated dataset of valid-invalid pairs. By testing on our constructed benchmark of twelve scenarios spanning over four physics domains, we show that our evaluation metric, Plausibility Preference Error (PPE), demonstrates strong alignment with human preference, outperforming state-of-the-art evaluator baselines. We then systematically benchmark intuitive physics understanding in current video diffusion models. Our study further analyses how model design and inference settings affect intuitive physics understanding and highlights domain-specific capacity variations across physical laws. Empirical results show that, despite current models struggling with complex and chaotic dynamics, there is a clear trend of improvement in physics understanding as model capacity and inference settings scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。