用动态3D几何体重建会随时间变化的场景,支持回放所有历史视角。
4D Primitive-Mâché: Glueing Primitives for Persistent 4D Scene Reconstruction
- 将场景拆解为运动的刚性3D几何体,通过2D对应关系联合优化其运动轨迹。
- 实现4D时空连续重建,在多物体数据集上显著优于现有方法。
- 适合需要物体持久性与时间回放能力的研究者,如机器人感知、数字孪生。
我们提出一种动态重建系统,以普通单目RGB视频为输入,输出完整且持续的场景重建结果。该系统不仅重建当前可见部分,还保留所有先前观测到的部分,从而可在任意时间点回放完整的4D场景。方法将场景分解为一组在空间中移动的刚性3D基本体,利用估计的稠密2D对应关系,通过优化流程联合推断这些基本体的刚性运动,得到4D动态重建——即随时间变化的三维几何结构。为此,我们引入运动外推机制,借助运动分组技术维持不可见物体的连续性。最终系统实现了4D时空感知,支持可回放的关节物体3D重建、多对象扫描及物体恒存性。在物体扫描和多对象数据集上,本方法在定量和定性指标上均显著优于现有方法。
原文摘要 · Abstract (English)
We present a dynamic reconstruction system that receives a casual monocular RGB video as input, and outputs a complete and persistent reconstruction of the scene. In other words, we reconstruct not only the the currently visible parts of the scene, but also all previously viewed parts, which enables replaying the complete reconstruction across all timesteps. Our method decomposes the scene into a set of rigid 3D primitives, which are assumed to be moving throughout the scene. Using estimated dense 2D correspondences, we jointly infer the rigid motion of these primitives through an optimisation pipeline, yielding a 4D reconstruction of the scene, i.e. providing 3D geometry dynamically moving through time. To achieve this, we also introduce a mechanism to extrapolate motion for objects that become invisible, employing motion-grouping techniques to maintain continuity. The resulting system enables 4D spatio-temporal awareness, offering capabilities such as replayable 3D reconstructions of articulated objects through time, multi-object scanning, and object permanence. On object scanning and multi-object datasets, our system significantly outperforms existing methods both quantitatively and qualitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。