视频扩散模型内部暗藏物理规律,无需显式训练即可捕捉真实运动逻辑。
The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show

- 通过反向积分速度场还原模型中间状态,探测其内在物理表征
- 在真实物理可判别任务上达81.27%准确率,超越专用表征学习模型
- 物理信息存在于去噪变压器内部,而非编码器输入,是生成过程的副产品
现代视频扩散模型生成的视频日益逼真且时间连贯,使其成为潜在的世界模拟器。然而,这些模型是否在内部编码了物理结构,还是仅复现训练中见过的运动模式,尚不明确。我们通过沿着已知物理合理性的真实视频潜轨迹探测视频扩散模型来研究这一问题。为获得此类轨迹,我们通过反向积分学习到的速度场,从干净视频潜变量回溯至噪声,从而获取模型的中间状态和注意力图。利用这些恢复的轨迹,我们发现物理合理性可在扩散变换器状态中线性解码,覆盖IntPhys和InfLevel数据集,平均准确率达81.27%,优于V-JEPA和VideoMAE等专用表征学习基线。令人惊讶的是,该信号在VAE潜变量输入中缺失,而是在去噪变压器内部出现,尽管模型未接受自监督预测目标训练。这些发现表明,物理有意义的表征可作为生成去噪过程的副产品自然涌现。
原文摘要 · Abstract (English)
Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators. Yet it remains unclear whether these models internally encode physical structure, or merely reproduce motion patterns seen during training. We study this question by probing video diffusion models along latent trajectories corresponding to real videos with known physical plausibility. To obtain such trajectories, we approximately invert the deterministic sampling process by integrating the learned velocity field backward from a clean video latent to noise, giving access to the model's intermediate states and attention maps. Using these recovered trajectories, we show that physical plausibility is linearly decodable from diffusion transformer states across IntPhys and InfLevel, reaching around 81.27% average accuracy and outperforming dedicated representation-learning baselines such as V-JEPA and VideoMAE. Surprisingly, this signal is absent from the VAE latent input and emerges inside the denoising transformer itself, despite the model not being trained with a self-supervised predictive objective. These findings suggest that physically meaningful representations can arise as a byproduct of generative denoising.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。