arXiv:2608.31025cs.CV2026-08

用物理约束的中间表示,从单目视频快速准确推断物体动力学。

Analytic Dynamics: Learning Physics-Grounded Representation for Fast Intrinsic Dynamics Inference from Monocular Videos

论文配图:Analytic Dynamics: Learning Physics-Grounded Representation for Fast Intrinsic Dynamics Inference from Monocular Videos
图 1 · 摘自论文原文
  • 引入物理可观测状态作为中间表示,连接视觉与内在动力学。
  • 在真实数据集上实现90%以上的材料参数回归准确率。
  • 适合需要高效物理推理的机器人、仿真系统应用。

从视觉观测中推断物体动力学对智能体理解与交互物理世界至关重要,但视觉证据与内在动力学之间存在根本性鸿沟。现有方法或依赖高成本的每场景优化,影响效率与可扩展性;或直接将视觉映射到动力学,缺乏中间物理抽象,易受外观与几何捷径干扰。为此,我们提出 Analytic Dynamics,一种前馈式动力学推断框架,引入一个介于视觉观测与内在动力学之间的物理基底动力学表示。具体而言,利用模拟中可用的特权物理状态(包括位置、位移和变形梯度场),学习结构化的动力学表示,该表示仅从视觉中难以发现。通过将视觉表征对齐至该空间,赋予视觉模型物理引导的归纳偏置,使其捕捉与动力学相关模式,用于材料模型分类与参数回归。为推动研究,我们构建了一个包含配对物理状态轨迹、渲染视频及真实材料模型与参数的动力学数据生成管道与基准。大量实验表明,Analytic Dynamics 能够从单目视频中实现高效、准确且泛化能力强的动力学推断。

原文摘要 · Abstract (English)

Inferring object dynamics from visual observations is essential for intelligent agents to reason about and interact with the physical world, yet remains challenging due to the fundamental gap between visual evidence and intrinsic dynamics. Existing methods either rely on costly per-scene optimization, limiting efficiency and scalability, or directly map visual evidence to intrinsic dynamics without intermediate physical abstractions, making them prone to appearance and geometry shortcuts. To bridge this gap, we propose Analytic Dynamics, a feed-forward dynamics inference framework that introduces an intermediate physics-grounded dynamics representation between visual observations and intrinsic dynamics. Specifically, we leverage privileged physical states, including position, displacement, and deformation gradient fields, which are available in simulation, to learn a structured dynamics representation that is difficult to discover from visual observations alone. By aligning visual representations with this space, we equip visual models with a physics-grounded inductive bias, guiding them to capture dynamics-relevant patterns for material model classification and parameter regression. To facilitate this research, we develop a dynamics data generation pipeline and benchmark containing paired physical state trajectories, rendered videos, and ground-truth material models and parameters. Extensive experiments demonstrate that Analytic Dynamics achieves efficient, accurate, and generalizable dynamics inference from monocular videos.

动力学推断物理建模单目视频神经表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。