视频模型生成的物体下落太慢,连伽利略定律都不懂,但少量数据就能改进。
Objects in Generated Videos Are Slower Than They Appear: Models Suffer Sub-Earth Gravity and Don't Know Galileo's Principle...for now
- 设计无单位对比实验,排除尺度干扰,精准测试物理规律
- 原始模型有效重力仅1.81米/秒²,远低于地球重力
- 仅用100段单球视频微调,重力提升至6.43米/秒²,接近真实值
视频生成模型被越来越多地视为潜在的世界模型,需编码并理解物理规律。我们研究其对基本物理定律——重力的表征。未经过训练的视频生成模型始终生成物体以更慢的加速度下落。然而,这些物理测试常受模糊的度量尺度干扰。我们首先检验观测到的物理错误是否由此类歧义(如帧率误判)引起。结果表明,即使进行时间重缩放也无法消除高方差的重力误差。为严格隔离物理表征与混淆因素,我们引入一种无单位、双物体测试协议,检验时间比 $t_1^2/t_2^2 = h_1/h_2$,该关系独立于重力 $g$、焦距和尺度。此相对测试揭示了对伽利略等效原理的违背。随后我们证明,通过针对性专化可部分弥补这一物理差距:仅用100段单球视频微调的轻量级低秩适配器,使有效重力 $g_{\mathrm{eff}}$ 从 $1.81\,\mathrm{m/s^2}$ 提升至 $6.43\,\mathrm{m/s^2}$(达到地球重力的65%)。该适配器还能零样本泛化至双球下落和斜面场景,初步表明仅用少量数据即可纠正特定物理规律。
原文摘要 · Abstract (English)
Video generators are increasingly evaluated as potential world models, which requires them to encode and understand physical laws. We investigate their representation of a fundamental law: gravity. Out-of-the-box video generators consistently generate objects falling at an effectively slower acceleration. However, these physical tests are often confounded by ambiguous metric scale. We first investigate if observed physical errors are artifacts of these ambiguities (e.g., incorrect frame rate assumptions). We find that even temporal rescaling cannot correct the high-variance gravity artifacts. To rigorously isolate the underlying physical representation from these confounds, we introduce a unit-free, two-object protocol that tests the timing ratio $t_1^2/t_2^2 = h_1/h_2$, a relationship independent of $g$, focal length, and scale. This relative test reveals violations of Galileo's equivalence principle. We then demonstrate that this physical gap can be partially mitigated with targeted specialization. A lightweight low-rank adaptor fine-tuned on only 100 single-ball clips raises $g_{\mathrm{eff}}$ from $1.81\,\mathrm{m/s^2}$ to $6.43\,\mathrm{m/s^2}$ (reaching $65\%$ of terrestrial gravity). This specialist adaptor also generalizes zero-shot to two-ball drops and inclined planes, offering initial evidence that specific physical laws can be corrected with minimal data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。