arXiv:2508.09811cs.CVcs.AI2025-08ICCV被引 13

从多视角视频中直接学习3D点的物理运动规律,无需人工标注。

TRACE: Learning 3D Gaussian Physical Dynamics from Multi-view Videos

  • 将每个3D点视为带大小和朝向的刚体粒子,直接学习其平移旋转动力学。
  • 在三个真实数据集和一个新合成数据集上,未来帧预测性能显著优于基线。
  • 通过物理参数聚类可自然分割多物体,适合需要动态理解与分割的任务。

本文旨在仅从动态多视角视频中建模3D场景的几何、外观及物理信息,无需任何人工标注。现有方法通常依赖物理约束损失或在神经网络中嵌入简单物理模型,但难以学习复杂运动物理,或需额外标签如物体类型、掩码等。我们提出新框架TRACE,用于建模复杂动态3D场景的运动物理。其核心创新在于:将每个3D点建模为具有空间尺寸与朝向的刚体粒子,直接学习每个粒子的平移-旋转动力学系统,并显式估计完整物理参数以控制其随时间的运动。在三个现有动态数据集和一个新构建的挑战性合成数据集上的大量实验表明,该方法在未来帧外推任务中表现卓越。该框架的一个优良特性是:可通过学习到的物理参数聚类,轻松实现多个物体或部件的分割。

原文摘要 · Abstract (English)

In this paper, we aim to model 3D scene geometry, appearance, and physical information just from dynamic multi-view videos in the absence of any human labels. By leveraging physics-informed losses as soft constraints or integrating simple physics models into neural nets, existing works often fail to learn complex motion physics, or doing so requires additional labels such as object types or masks. We propose a new framework named TRACE to model the motion physics of complex dynamic 3D scenes. The key novelty of our method is that, by formulating each 3D point as a rigid particle with size and orientation in space, we directly learn a translation rotation dynamics system for each particle, explicitly estimating a complete set of physical parameters to govern the particle's motion over time. Extensive experiments on three existing dynamic datasets and one newly created challenging synthetic datasets demonstrate the extraordinary performance of our method over baselines in the task of future frame extrapolation. A nice property of our framework is that multiple objects or parts can be easily segmented just by clustering the learned physical parameters.

3D重建物理建模动态场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。