从单目视频推断物体物理动态,支持用户实时施力交互。
LaGSplat: Inferring Physics-Governed Interactive Simulation from Monocular Video Using Latent Lagrangian Gaussian Splatting

- 用隐空间坐标同时表征物理系统状态与渲染点位置。
- 可对未见过的外力实时响应,3D/2D画面即时生成。
- 适合需要物理真实感的交互式模拟场景。
我们提出LaGSplat(隐空间拉格朗日高斯点云),一种从一个或少量单目视频中推断可交互、受物理规律约束的动力学框架。推理时,用户可对拍摄对象(刚体或变形体)施加训练中未测量、标注或出现过的外部力。这得益于低维隐状态 $\mathbf{q} \in \mathbb{R}^d$ 的双重角色:既是学习得到的耗散拉格朗日系统的广义坐标,又是高斯点云解码器的条件变量。该解码器具有归纳偏置,其基元为随物体移动的显式点 $μ_i(\mathbf{q})$,使图像中的力 $f$ 能反向映射为隐空间广义力 $J(\mathbf{q})^\top f$ 并进入运动方程——这是像素空间(CNN)或神经场(NeRF)解码器无法做到的。我们在从刚体到变形体、从自主系统到受迫真实系统的递进测试案例上验证了方法有效性,结合单目视频与传感器数据。进一步展示了交互能力:任意大小和方向的力可在任意时刻施加于重建对象,其响应可实时渲染为2D或3D图像。假设在少数广义坐标上存在耗散欧拉-拉格朗日方程,以牺牲通用性换取对未知力的有界且合理的响应,而无约束预测会发散。
原文摘要 · Abstract (English)
We present LaGSplat (Latent Lagrangian Gaussian Splatting), a framework that infers interactive, physics-governed dynamics from one or a few monocular videos. At inference it lets a user push on the filmed object, rigid or deformable, with an external force that was never measured, annotated, or seen during training. This is possible because a low-dimensional latent state $\mathbf{q} \in \mathbb{R}^d$ plays two roles at once: it is the generalised coordinate of a learned dissipative Lagrangian and the conditioning variable of a Gaussian Splatting decoder. The inductive bias of this decoder, whose primitives are explicit points $μ_i(\mathbf{q})$ that move with the object, is what lets a force $f$ applied in the image pull back into a latent generalised force $J(\mathbf{q})^\top f$ and enter the equations of motion, which pixel-space (CNN) or neural-field (NeRF) decoders cannot do. We validate LaGSplat on test cases of increasing difficulty, from rigid to deformable and from autonomous to forced real systems, combining monocular video and sensor measurements. We further demonstrate interactive use: forces of arbitrary magnitude and direction can be applied to the reconstructed object at any time, its response rendered in real time, in 2D or 3D. Assuming a dissipative Euler-Lagrange equation over a few generalised coordinates trades generality for a bounded, plausible response to unseen forces, where an unconstrained predictor diverges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。