arXiv:2608.22102cs.CVcs.AI2026-08

从单目视频中学习3D高斯表示的隐式物理规律,提升动态物体建模精度。

Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos

论文配图:Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos
图 1 · 摘自论文原文
  • 结合LoRA与两个对齐模块,从单视角视频中学习物理动态。
  • 合成数据上比最强基线降低48%的Chamfer Distance,鲁棒性更强。
  • 适合需要物理可解释性的动态3D重建任务,如真实场景建模。

我们提出GCA(Gaussian Constitutive Alignment),一种从由3D高斯表示的可变形物体单目动态视频中学习隐式本构规律的框架。在静态多视角扫描用于几何初始化的前提下,该方法仅需一个固定视角的运动物体视频即可学习内在物理动态。现有隐式方法在噪声监督下易陷入局部最优且缺乏物理可解释性,而显式方法依赖预设本构方程,泛化能力差且在单目设置下不稳定。为解决上述问题,我们的框架融合基于LoRA的适配与两个关键对齐模块:首先提出基于秩的深度-几何锚点(RDGA),通过尺度不变的秩对齐实现单目动态观测下的稳健几何约束,降低对不可靠像素级颜色监督的依赖;其次引入本构先验正则化器(CPR),将经典本构模型作为软可微先验,规范优化过程,同时保持隐式建模的灵活性——即使实际材料未包含在假设中也有效。在合成、真实到仿真及真实世界数据集上的大量实验表明,GCA优于现有方法,在合成基准上比最强基线降低48%的Chamfer Distance,且在单目监督下仍具鲁棒性。

原文摘要 · Abstract (English)

We present GCA (Gaussian Constitutive Alignment), a framework for learning implicit constitutive laws from monocular dynamic video of deformable objects represented by 3D Gaussians. Given a static multi-view scan for geometric initialization, our method learns intrinsic physical dynamics solely from a single fixed-viewpoint video of the moving object. Existing implicit methods often suffer from local minima under noisy supervision and lack physical interpretability, while explicit approaches rely on predefined constitutive equations, limiting generalizability and becoming unstable in monocular settings. To address these challenges, our framework unifies LoRA-based adaptation with two key alignment modules. First, we propose Rank-based Depth-Geometric Anchors (RDGA) to establish robust geometric constraints from monocular dynamic observations via scale-invariant rank-based depth alignment, reducing the reliance on unreliable pixel-level color supervision. Second, a Constitutive Prior Regularizer (CPR) integrates classical constitutive models as soft differentiable priors, regularizing the optimization while preserving the flexibility of implicit modeling---even when the actual material is absent from the hypotheses. Extensive experiments on synthetic, real-to-sim, and real-world datasets demonstrate that GCA outperforms existing methods, achieving 48% lower Chamfer Distance than the strongest baseline on synthetic benchmarks while remaining robust under monocular supervision.

3D重建隐式建模物理驱动单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。