揭秘视频模型如何隐式表征物理信息,发现关键的物理涌现区。
Interpreting Physics in Video World Models
- 通过层间探查与注意力消融,定位物理信息在编码器中的涌现位置。
- 速度加速度早期可读,方向仅在中间层出现,且呈高维环形结构。
- 模型用分布式表示而非分解变量,适合理解视觉物理推理机制。
视频世界模型在直观物理任务上表现优异,但其内部是否依赖物理变量的因子化表示仍不明确。本文首次对大规模视频编码器进行可解释性分析,结合层间探测、子空间几何、像素级解码与定向注意力消融,揭示物理信息在编码器中的组织方式。研究发现存在一个显著的中间深度转折点——物理涌现区(Physics Emergence Zone),在此处物理变量开始可被访问,随后迅速达到峰值并随输出层退化。分解运动后发现,速度、加速度从早期层即可读取,而运动方向仅在物理涌现区才显现。方向以高维群体结构编码,具有环形几何特征,需多特征协同干预才能操控。结果表明,现代视频模型未采用类物理引擎的因子化表示,而是通过分布式的隐式表征实现物理预测。
原文摘要 · Abstract (English)
A long-standing question in physical reasoning is whether video-based models need to rely on factorized representations of physical variables in order to make physically accurate predictions, or whether they can implicitly represent such variables in a task-specific, distributed manner. While modern video world models achieve strong performance on intuitive physics benchmarks, it remains unclear which of these representational regimes they implement internally. Here, we present the first interpretability study to directly examine physical representations inside large-scale video encoders. Using layerwise probing, subspace geometry, patch-level decoding, and targeted attention ablations, we characterize where physical information becomes accessible and how it is organized within encoder-based video transformers. Across architectures, we identify a sharp intermediate-depth transition -- which we call the Physics Emergence Zone -- at which physical variables become accessible. Physics-related representations peak shortly after this transition and degrade toward the output layers. Decomposing motion into explicit variables, we find that scalar quantities such as speed and acceleration are available from early layers onwards, whereas motion direction becomes accessible only at the Physics Emergence Zone. Notably, we find that direction is encoded through a high-dimensional population structure with circular geometry, requiring coordinated multi-feature intervention to control. These findings suggest that modern video models do not use factorized representations of physical variables like a classical physics engine. Instead, they use a distributed representation that is nonetheless sufficient for making physical predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。