首个联合实现非刚性表面追踪与重建的4D SLAM方法
4DTAM: Non-Rigid Tracking and Mapping via Dynamic Surface Gaussians
- 用可微渲染结合高斯表面原语建模动态场景
- 通过MLP变形场实现时序非刚性形变精准重建
- 构建首个公开合成4D数据集支持算法评估
我们提出首个联合实现相机定位与非刚性表面重建的4D追踪与建图方法,基于可微渲染从彩色图像流中融合深度测量或预测结果,同步优化场景几何、外观、动态特性及相机自运动。针对自然环境复杂非刚性运动带来的挑战,传统2.5D信号下优化空间维度高导致问题病态,我们引入基于高斯表面原语的SLAM方法,比3D高斯更有效利用深度信号,提升表面重建精度。为建模非刚性形变,采用多层感知机(MLP)表示形变场,并提出新型相机位姿估计方法及表面正则化项以支持时空重建。此外,4D SLAM研究受限于缺乏可靠真值与评估协议,主要因消费级传感器难以捕捉4D数据。为此,我们构建了一个包含多样化日常物体运动的开源合成数据集,依托大规模物体模型与动画建模技术。综上,本工作通过新方法与评估协议,推动现代4D-SLAM研究发展。
原文摘要 · Abstract (English)
We propose the first 4D tracking and mapping method that jointly performs camera localization and non-rigid surface reconstruction via differentiable rendering. Our approach captures 4D scenes from an online stream of color images with depth measurements or predictions by jointly optimizing scene geometry, appearance, dynamics, and camera ego-motion. Although natural environments exhibit complex non-rigid motions, 4D-SLAM remains relatively underexplored due to its inherent challenges; even with 2.5D signals, the problem is ill-posed because of the high dimensionality of the optimization space. To overcome these challenges, we first introduce a SLAM method based on Gaussian surface primitives that leverages depth signals more effectively than 3D Gaussians, thereby achieving accurate surface reconstruction. To further model non-rigid deformations, we employ a warp-field represented by a multi-layer perceptron (MLP) and introduce a novel camera pose estimation technique along with surface regularization terms that facilitate spatio-temporal reconstruction. In addition to these algorithmic challenges, a significant hurdle in 4D SLAM research is the lack of reliable ground truth and evaluation protocols, primarily due to the difficulty of 4D capture using commodity sensors. To address this, we present a novel open synthetic dataset of everyday objects with diverse motions, leveraging large-scale object models and animation modeling. In summary, we open up the modern 4D-SLAM research by introducing a novel method and evaluation protocols grounded in modern vision and rendering techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。