从视觉观察中推断可解释的物体物理规律,提升仿真真实性和泛化能力。
VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization
- 通过双层优化框架,用大模型生成并迭代改进物理定律
- 在合成与真实数据上均显著优于现有方法,泛化性能强
- 适合需要可解释物理仿真的研究者和工程师
物体的内在动力学决定了其在现实世界中的物理行为,在实现与3D资产的物理合理交互仿真中起关键作用。现有方法从视觉观测中推断内在动力学,面临两大挑战:一类依赖人工定义的本构先验,难以对齐真实动力学;另一类使用神经网络建模,可解释性差且泛化能力弱。为此,我们提出VisionLaw,一种基于双层优化的框架,从视觉观测中推断可解释的内在动力学表达。上层采用大模型驱动的解耦本构演化策略,让大模型作为物理专家生成并修正本构定律,内置解耦机制显著降低搜索复杂度;下层引入视觉引导的本构评估机制,利用视觉仿真评估生成定律与底层内在动力学的一致性,从而指导上层演化。在合成与真实世界数据集上的实验表明,VisionLaw能有效从视觉观测中推断出可解释的内在动力学,显著优于现有最先进方法,并在新场景的交互仿真中表现出强泛化能力。
原文摘要 · Abstract (English)
The intrinsic dynamics of an object governs its physical behavior in the real world, playing a critical role in enabling physically plausible interactive simulation with 3D assets. Existing methods have attempted to infer the intrinsic dynamics of objects from visual observations, but generally face two major challenges: one line of work relies on manually defined constitutive priors, making it difficult to align with actual intrinsic dynamics; the other models intrinsic dynamics using neural networks, resulting in limited interpretability and poor generalization. To address these challenges, we propose VisionLaw, a bilevel optimization framework that infers interpretable expressions of intrinsic dynamics from visual observations. At the upper level, we introduce an LLMs-driven decoupled constitutive evolution strategy, where LLMs are prompted to act as physics experts to generate and revise constitutive laws, with a built-in decoupling mechanism that substantially reduces the search complexity of LLMs. At the lower level, we introduce a vision-guided constitutive evaluation mechanism, which utilizes visual simulation to evaluate the consistency between the generated constitutive law and the underlying intrinsic dynamics, thereby guiding the upper-level evolution. Experiments on both synthetic and real-world datasets demonstrate that VisionLaw can effectively infer interpretable intrinsic dynamics from visual observations. It significantly outperforms existing state-of-the-art methods and exhibits strong generalization for interactive simulation in novel scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。