从视频中精准识别物理规律,仅需少量轨迹即可恢复真实参数。
Physics from Video: Identifiability of Time-Invariant Second-Order ODEs under Minimal Trajectory Conditions

- 通过水平集斜率覆盖条件,保证潜在空间与真实状态局部仿射。
- 欠阻尼系统仅需一个视频片段即可识别,其他情形需三段不同轨迹。
- 无需像素重建,用方差下界正则化防止潜在表示坍塌,适合物理建模研究者。
弥合视觉真实与物理理解之间的鸿沟是基于视频的世界模型的核心挑战。本文研究从原始像素中识别连续时间物理规律的结构可辨识性,重点关注仅使用编码器的流水线能否唯一恢复二阶线性常微分方程的参数。我们证明,水平集斜率覆盖率条件可确保学习到的潜在空间在局部上仿射于真实物理状态,从而实现参数精确恢复。理论首次刻画了跨阻尼区间的最小数据需求:欠阻尼系统仅需单个视频片段即可辨识,而其他阻尼类型则需三个不同轨迹。此外,我们引入方差下界正则化以稳定无解码器目标,防止潜在表示坍塌。在合成与真实数据上的验证表明,该方法可从视频中可靠估计可解释的物理常数,无需高计算量的像素重建,兼顾物理正确性与透明性。代码已开源:https://github.com/wenjiewang3/PhysicsFromVideo。
原文摘要 · Abstract (English)
Bridging the gap between visual realism and physical understanding is a core challenge for video-based world models. We study the structural identifiability of continuous-time physical laws from raw pixels, focusing on whether an encoder-only pipeline can uniquely recover the parameters of second-order linear ODEs. We prove that a level-set slope-coverage condition ensures the learned latent space is locally affine to the true physical state, enabling exact parameter recovery. Our theory provides the first characterization of minimal data requirements across damping regimes, establishing that underdamped systems are identifiable from a single video clip, whereas other regimes require three diverse trajectories. We further introduce a variance-floor regularizer to stabilize the decoder-free objective and prevent latent collapse. Validated on synthetic and real-world data, our approach demonstrates that interpretable physical constants can be reliably estimated from video without the need for compute-intensive pixel reconstruction, ensuring both physical correctness and transparency. Code is available at https://github.com/wenjiewang3/PhysicsFromVideo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。