arXiv:2605.10586cs.CV2026-05中稿 · ed

从多视角视频中学习3D动态场景的物理因果关系,无需人工标注。

CausalGS: Learning Physical Causality of 3D Dynamic Scenes with Gaussian Representations

论文配图:CausalGS: Learning Physical Causality of 3D Dynamic Scenes with Gaussian Representations
图 1 · 摘自论文原文
  • 通过逆物理推断分离初始速度场与材料属性,解耦复杂动态问题。
  • 在长时序未来帧预测任务上超越现有方法,实现高精度插值。
  • 仅凭视觉信息即可捕捉多物理量间的因果关系,适合视觉建模研究者。

从视频数据中学习能理解物理规律并预测物体未来轨迹的物理模型是人工智能中的重大挑战。以往方法或使用偏微分方程作为软约束(如PINN损失),或整合物理模拟器到神经网络中,但通常依赖强先验或高质量几何重建。本文提出CausalGS框架,仅从多视角视频中学习复杂动态3D场景的因果动力学,且无需显式先验。其核心是一个逆物理推断模块,将复杂的动态问题分解为两个因子的联合推断:表征场景运动学的初始速度场,以及控制其动态的内在材料属性。该推断得到的物理信息被用于可微分物理模拟器中,以物理正则化方式指导学习过程。大量实验表明,CausalGS在极具挑战性的长期未来帧外推任务中超越当前最优方法,同时在新视角插值任务中也表现优异。关键在于,无需任何人工标注,模型仅从视觉观测中即能学习多个物理属性间的复杂交互,并理解驱动场景动态演化的因果关系。

原文摘要 · Abstract (English)

Learning a physical model from video data that can comprehend physical laws and predict the future trajectories of objects is a formidable challenge in artificial intelligence. Prior approaches either leverage various Partial Differential Equations (PDEs) as soft constraints in the form of PINN losses, or integrate physics simulators into neural networks; however, they often rely on strong priors or high-quality geometry reconstruction. In this paper, we propose CausalGS, a framework that learns the causal dynamics of complex dynamic 3D scenes solely from multi-view videos, while dispensing with the reliance on explicit priors. At its core is an inverse physics inference module that decouples the complex dynamics problem from the video into the joint inference of two factors: the initial velocity field representing the scene's kinematics, and the intrinsic material properties governing its dynamics. This inferred physical information is then utilized within a differentiable physics simulator to guide the learning process in a physics-regularized manner. Extensive experiments demonstrate that CausalGS surpasses the state-of-the-art on the highly challenging task of long-term future frame extrapolation, while also exhibiting advanced performance in novel view interpolation. Crucially, our work shows that, without any human annotation, the model is able to learn the complex interactions between multiple physical properties and understand the causal relationships driving the scene's dynamic evolution, solely from visual observations.

3D动态建模因果推断物理模拟多视角视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。