arXiv:2410.13862cs.CV2024-10CVPR被引 299

将高斯点阵与深度估计结合,实现快速高质量三维重建。

DepthSplat: Connecting Gaussian Splatting and Depth

  • 用预训练单目深度特征构建鲁棒多视角深度模型。
  • 在多个数据集上同时提升深度估计与新视角合成效果。
  • 12张图输入仅需0.6秒完成前向重建,适合实时应用。

高斯点阵与单视图深度估计通常被独立研究。本文提出DepthSplat,连接二者并探究其交互作用。我们首先利用预训练单目深度特征,构建鲁棒的多视角深度模型,从而实现高质量前向3D高斯点阵重建。同时表明,高斯点阵可作为无监督预训练目标,从大规模多视角姿态数据集中学习强大的深度模型。通过大量消融实验与跨任务迁移验证了两者的协同效应。DepthSplat在ScanNet、RealEstate10K和DL3DV数据集上均达到当前最优的深度估计与新视角合成性能,展现了两任务间的相互增益。此外,该方法可在0.6秒内完成12张输入图像(512x960分辨率)的前向重建。

原文摘要 · Abstract (English)

Gaussian splatting and single-view depth estimation are typically studied in isolation. In this paper, we present DepthSplat to connect Gaussian splatting and depth estimation and study their interactions. More specifically, we first contribute a robust multi-view depth model by leveraging pre-trained monocular depth features, leading to high-quality feed-forward 3D Gaussian splatting reconstructions. We also show that Gaussian splatting can serve as an unsupervised pre-training objective for learning powerful depth models from large-scale multi-view posed datasets. We validate the synergy between Gaussian splatting and depth estimation through extensive ablation and cross-task transfer experiments. Our DepthSplat achieves state-of-the-art performance on ScanNet, RealEstate10K and DL3DV datasets in terms of both depth estimation and novel view synthesis, demonstrating the mutual benefits of connecting both tasks. In addition, DepthSplat enables feed-forward reconstruction from 12 input views (512x960 resolutions) in 0.6 seconds.

三维重建深度估计高斯点阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。