无需物体先验,从视频中学习复杂3D物理运动。
FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity
- 用高斯速度场建模动态场景物理,避免依赖微分方程损失。
- 在三个公开数据集和新采集的真实数据上实现更优未来帧预测。
- 学习到的物理编码可捕捉真实3D运动模式,无需人工标注。
本文旨在仅从多视角视频中建模3D场景的几何、外观及底层物理。现有方法常依赖各类控制偏微分方程(PDE)作为物理约束损失,或需物体掩码、类别等先验,难以处理边界复杂运动。为此,本文提出FreeGave,无需任何物体先验即可学习复杂动态3D场景的物理规律。核心是引入物理编码,并设计散度为零模块,直接估计每个高斯点的速度场,规避低效的物理残差损失。在三个公开数据集及一个新收集的挑战性真实数据集上的实验表明,该方法在未来帧外推和运动分割任务中表现优异。尤其值得注意的是,对学习到的物理编码分析发现,其能有效捕捉无标签训练下的真实3D运动模式。
原文摘要 · Abstract (English)
In this paper, we aim to model 3D scene geometry, appearance, and the underlying physics purely from multi-view videos. By applying various governing PDEs as PINN losses or incorporating physics simulation into neural networks, existing works often fail to learn complex physical motions at boundaries or require object priors such as masks or types. In this paper, we propose FreeGave to learn the physics of complex dynamic 3D scenes without needing any object priors. The key to our approach is to introduce a physics code followed by a carefully designed divergence-free module for estimating a per-Gaussian velocity field, without relying on the inefficient PINN losses. Extensive experiments on three public datasets and a newly collected challenging real-world dataset demonstrate the superior performance of our method for future frame extrapolation and motion segmentation. Most notably, our investigation into the learned physics codes reveals that they truly learn meaningful 3D physical motion patterns in the absence of any human labels in training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。