统一时空对齐提升稀疏视角4D重建精度
UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment

- 将时空深度对齐整合为统一框架,无需分割掩码
- 在Ego-Exo4D和EgoHuman上实现更优的几何一致性
- 适合动态场景下新视角/新时间合成任务
本文针对稀疏视角视频的4D重建难题,提出统一时空深度对齐框架。传统方法分阶段处理空间与时间对齐,依赖前景分割掩码且忽略时序线索。本方法将多视角、多时间的深度图建模为一组时空神经场,隐式捕捉深度图间的时空相关性,无需外部分割或追踪模型,加速收敛。同时引入多视图深度顺序损失,结合尺度平移不变损失,显著提升深度质量。对齐后的深度用于初始化并监督高斯点云渲染模型,实现4D重建。在Ego-Exo4D和EgoHuman数据集上的实验表明,该方法大幅提升了基于动态高斯溅射的重建性能,在新时间/视角合成及几何精度与一致性方面表现优异。
原文摘要 · Abstract (English)
In this paper, we address the challenging problem of 4D reconstruction from sparse-view videos. This setup usually relies on monocular depth estimation to provide priors for the reconstruction model. A key challenge arises from limited cross-view overlap and temporal variation, making monocular depth predictions inconsistent across views and time. Existing methods align spatial and temporal dimensions in separate stages, requiring foreground segmentation masks while failing to leverage temporal cues for cross-view alignment. Contrary to these methods, we propose a unified spatial-temporal depth alignment framework that jointly resolves cross-view and cross-time inconsistencies without distinguishing foreground/background. Our method represents depth maps across views and time as a set of spatio-temporal neural fields. This representation not only yields fast convergence, but also captures spatio-temporal correlation among depth maps implicitly, without dependence on external segmentation/tracking models. We also propose a multi-view depth-order loss while leveraging the classic scale-and-shift-invariant loss to further improve the final depth quality. The aligned depths initialize and supervise Gaussian splatting models for 4D reconstruction. Experiments on Ego-Exo4D and EgoHuman demonstrate that our improved depth alignment substantially benefits dynamic Gaussian-splatting-based reconstruction methods for novel-time/view synthesis and geometry accuracy/consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。