视频生成模型难以捕捉物理规律,仅靠数据无法实现真正泛化。
How Far is Video Generation from World Model: A Physical Law Perspective

- 构建2D物理模拟环境,用经典力学生成可控视频数据
- 模型在分布内完美泛化,但分布外场景完全失效
- 模型依赖记忆而非抽象规则,优先考虑颜色等表面特征
OpenAI的Sora展示了视频生成在构建符合基本物理定律的世界模型方面的潜力。然而,仅从视觉数据中不依赖人类先验就能发现这些定律的能力值得怀疑。一个真正学习物理规律的世界模型应能对细微变化保持鲁棒性,并在未见场景中正确外推。本文在三种关键场景下评估:分布内、分布外和组合泛化。我们构建了一个2D模拟测试平台,通过经典力学定律确定性地生成物体运动与碰撞视频,提供大规模实验所需的数据,并可量化评估生成视频是否符合物理规律。训练基于扩散模型的视频生成模型根据初始帧预测物体运动。规模实验显示:在分布内实现完美泛化,组合泛化表现出可测量的缩放行为,但在分布外场景完全失败。进一步实验揭示两个关键洞察:(1) 模型未能抽象出通用物理规则,而是表现出‘案例式’泛化,即模仿最接近的训练样本;(2) 在泛化到新情况时,模型参考训练数据时优先级为:颜色 > 大小 > 速度 > 形状。研究表明,单纯扩大规模不足以使视频生成模型发现基础物理规律,尽管这在Sora的成功中起到了作用。
原文摘要 · Abstract (English)
OpenAI's Sora highlights the potential of video generation for developing world models that adhere to fundamental physical laws. However, the ability of video generation models to discover such laws purely from visual data without human priors can be questioned. A world model learning the true law should give predictions robust to nuances and correctly extrapolate on unseen scenarios. In this work, we evaluate across three key scenarios: in-distribution, out-of-distribution, and combinatorial generalization. We developed a 2D simulation testbed for object movement and collisions to generate videos deterministically governed by one or more classical mechanics laws. This provides an unlimited supply of data for large-scale experimentation and enables quantitative evaluation of whether the generated videos adhere to physical laws. We trained diffusion-based video generation models to predict object movements based on initial frames. Our scaling experiments show perfect generalization within the distribution, measurable scaling behavior for combinatorial generalization, but failure in out-of-distribution scenarios. Further experiments reveal two key insights about the generalization mechanisms of these models: (1) the models fail to abstract general physical rules and instead exhibit "case-based" generalization behavior, i.e., mimicking the closest training example; (2) when generalizing to new cases, models are observed to prioritize different factors when referencing training data: color > size > velocity > shape. Our study suggests that scaling alone is insufficient for video generation models to uncover fundamental physical laws, despite its role in Sora's broader success. See our project page at https://phyworld.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。