统一时空融合框架,实现动态驾驶场景高保真重建
UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction
- 构建3D潜在骨架,融合多视角时序信息
- 新视图合成性能领先,覆盖原视野外区域
- 适合自动驾驶场景建模与实时渲染
面向自动驾驶的前馈式3D重建已取得快速进展,但现有方法在稀疏、非重叠摄像头视角和复杂场景动态性方面仍面临挑战。我们提出UniSplat,一种通用前馈框架,通过统一的潜在时空融合实现鲁棒的动态场景重建。UniSplat构建一个3D潜在骨架,利用预训练基础模型捕捉几何与语义场景上下文。为高效融合空间视角与时间帧信息,引入直接在3D骨架内操作的高效融合机制,实现一致的时空对齐。为确保完整且细节丰富的重建,设计双分支解码器,结合点锚定优化与体素生成,从融合骨架中生成动态感知高斯点;并维护静态高斯点的持久记忆,支持超出当前摄像头覆盖范围的流式场景补全。在真实世界数据集上的大量实验表明,UniSplat在新视图合成上达到当前最佳性能,即使在原始摄像头覆盖范围外也能提供鲁棒且高质量的渲染结果。
原文摘要 · Abstract (English)
Feed-forward 3D reconstruction for autonomous driving has advanced rapidly, yet existing methods struggle with the joint challenges of sparse, non-overlapping camera views and complex scene dynamics. We present UniSplat, a general feed-forward framework that learns robust dynamic scene reconstruction through unified latent spatio-temporal fusion. UniSplat constructs a 3D latent scaffold, a structured representation that captures geometric and semantic scene context by leveraging pretrained foundation models. To effectively integrate information across spatial views and temporal frames, we introduce an efficient fusion mechanism that operates directly within the 3D scaffold, enabling consistent spatio-temporal alignment. To ensure complete and detailed reconstructions, we design a dual-branch decoder that generates dynamic-aware Gaussians from the fused scaffold by combining point-anchored refinement with voxel-based generation, and maintain a persistent memory of static Gaussians to enable streaming scene completion beyond current camera coverage. Extensive experiments on real-world datasets demonstrate that UniSplat achieves state-of-the-art performance in novel view synthesis, while providing robust and high-quality renderings even for viewpoints outside the original camera coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。