测试视频模型是否真懂物理,发现画得再像也不代表理解物理规律。
Do generative video models understand physical principles?
- 构建物理智商评测集,要求模型理解流体、光学等真实物理原理。
- 现有模型在物理理解上普遍薄弱,与画面逼真度无关。
- 部分问题可解决,说明仅靠观察学物理有潜力但难度仍大。
AI视频生成正经历革命性进展,质量与真实感迅速提升。这引发了科学界热议:视频模型是否学习了‘世界模型’并发现物理规律?抑或只是高级像素预测器,虽视觉逼真却无真实物理理解?为此,我们开发了Physics-IQ——一个综合性基准数据集,仅凭对流体力学、光学、固体力学、磁学和热力学等物理原理的深层理解才能解答。测试涵盖Sora、Runway、Pika、Lumiere、Stable Video Diffusion和VideoPoet等主流模型,结果显示其物理理解严重受限,且与视觉真实感无关。尽管部分测试案例已可成功解决,表明仅通过观察获取某些物理原则可能可行,但挑战依然巨大。我们预计未来将快速进步,但本研究证明:视觉真实不等于物理理解。项目主页:https://physics-iq.github.io;代码:https://github.com/google-deepmind/physics-IQ-benchmark。
原文摘要 · Abstract (English)
AI video generation is undergoing a revolution, with quality and realism advancing rapidly. These advances have led to a passionate scientific debate: Do video models learn "world models" that discover laws of physics -- or, alternatively, are they merely sophisticated pixel predictors that achieve visual realism without understanding the physical principles of reality? We address this question by developing Physics-IQ, a comprehensive benchmark dataset that can only be solved by acquiring a deep understanding of various physical principles, like fluid dynamics, optics, solid mechanics, magnetism and thermodynamics. We find that across a range of current models (Sora, Runway, Pika, Lumiere, Stable Video Diffusion, and VideoPoet), physical understanding is severely limited, and unrelated to visual realism. At the same time, some test cases can already be successfully solved. This indicates that acquiring certain physical principles from observation alone may be possible, but significant challenges remain. While we expect rapid advances ahead, our work demonstrates that visual realism does not imply physical understanding. Our project page is at https://physics-iq.github.io; code at https://github.com/google-deepmind/physics-IQ-benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。