arXiv:2505.02175cs.CV2025-05被引 9

用2D高斯点云实现快速通用的稀疏视角三维重建与新视角合成。

SparSplat: Fast Multi-View Reconstruction with Generalizable 2D Gaussian Splatting

  • 基于多视角学习框架,前馈式回归2D高斯表面参数进行三维重建。
  • 在DTU数据集上达到最优的切比雪夫距离,新视角合成质量领先。
  • 推理速度比隐式表示方法快近100倍,适用于真实场景部署。

从多视角图像中恢复三维信息(多视角立体重建与新视角合成)在稀疏视角场景下极具挑战性。3D高斯点云(3DGS)实现了实时、逼真的新视角合成。随后,2D高斯点云(2DGS)通过透视精确的2D高斯基元光栅化,提升了渲染中的几何表示精度,同时保持实时性能。近期方法在通用多视角学习框架下,利用3DGS实现稀疏实时新视角合成,以回归3D高斯参数。本文进一步拓展该方向,联合解决可泛化的稀疏三维重建与新视角合成问题。提出一种基于多视角学习的前馈式管道,直接回归2DGS表面元素参数,实现稀疏视图下的三维形状重建与新视角合成。结果表明,该模型可受益于预训练的多视角深度视觉特征,在DTU稀疏重建基准上以切比雪夫距离达到当前最优,新视角合成表现同样领先。在BlendedMVS与Tanks and Temples数据集上也展现出强泛化能力。相比基于隐式表示体素渲染的前馈式稀疏重建方法,本模型不仅性能更优,且推理速度提升近两个数量级。

原文摘要 · Abstract (English)

Recovering 3D information from scenes via multi-view stereo reconstruction (MVS) and novel view synthesis (NVS) is inherently challenging, particularly in scenarios involving sparse-view setups. The advent of 3D Gaussian Splatting (3DGS) enabled real-time, photorealistic NVS. Following this, 2D Gaussian Splatting (2DGS) leveraged perspective accurate 2D Gaussian primitive rasterization to achieve accurate geometry representation during rendering, improving 3D scene reconstruction while maintaining real-time performance. Recent approaches have tackled the problem of sparse real-time NVS using 3DGS within a generalizable, MVS-based learning framework to regress 3D Gaussian parameters. Our work extends this line of research by addressing the challenge of generalizable sparse 3D reconstruction and NVS jointly, and manages to perform successfully at both tasks. We propose an MVS-based learning pipeline that regresses 2DGS surface element parameters in a feed-forward fashion to perform 3D shape reconstruction and NVS from sparse-view images. We further show that our generalizable pipeline can benefit from preexisting foundational multi-view deep visual features. The resulting model attains the state-of-the-art results on the DTU sparse 3D reconstruction benchmark in terms of Chamfer distance to ground-truth, as-well as state-of-the-art NVS. It also demonstrates strong generalization on the BlendedMVS and Tanks and Temples datasets. We note that our model outperforms the prior state-of-the-art in feed-forward sparse view reconstruction based on volume rendering of implicit representations, while offering an almost 2 orders of magnitude higher inference speed.

3D重建2D高斯新视角合成实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。