arXiv:2411.06390cs.CV2024-11ICLR被引 17

提出首个专用于3D高斯点云的点变换器,提升极端新视角渲染质量。

SplatFormer: Point Transformer for Robust 3D Gaussian Splatting

  • 设计点变换器SplatFormer,直接作用于3D高斯点云进行单次前向优化
  • 在极端新视角下显著提升渲染质量,超越现有正则化与多场景模型
  • 适合沉浸式自由视角渲染、导航等需要鲁棒泛化的应用

3D高斯喷溅(3DGS)最近实现了逼真的三维重建,具备高视觉保真度和实时性能。然而,当测试视角偏离训练时的相机角度时,渲染质量显著下降,严重制约了其在沉浸式自由视角渲染与导航中的应用。本文对3DGS及相关新视图合成方法在分布外(OOD)相机场景下的表现进行了全面评估。通过构建包含合成与真实数据集的多样化测试用例,我们发现多数现有方法,包括采用各类正则化技术和数据驱动先验的方法,均难以有效泛化至OOD视图。为此,我们提出SplatFormer,首个专为高斯喷溅设计的点变换器模型。SplatFormer接收仅在有限训练视角下优化的初始3DGS集合,在单次前向传播中对其进行精炼,有效消除分布外测试视图中的潜在伪影。据我们所知,这是首次成功将点变换器直接应用于3DGS集合,突破了此前多场景训练方法在推理时仅能处理有限输入视角的限制。该模型在极端新视角下显著提升渲染质量,达到当前最优性能,优于多种3DGS正则化技术、专为稀疏视图合成设计的多场景模型以及基于扩散的框架。

原文摘要 · Abstract (English)

3D Gaussian Splatting (3DGS) has recently transformed photorealistic reconstruction, achieving high visual fidelity and real-time performance. However, rendering quality significantly deteriorates when test views deviate from the camera angles used during training, posing a major challenge for applications in immersive free-viewpoint rendering and navigation. In this work, we conduct a comprehensive evaluation of 3DGS and related novel view synthesis methods under out-of-distribution (OOD) test camera scenarios. By creating diverse test cases with synthetic and real-world datasets, we demonstrate that most existing methods, including those incorporating various regularization techniques and data-driven priors, struggle to generalize effectively to OOD views. To address this limitation, we introduce SplatFormer, the first point transformer model specifically designed to operate on Gaussian splats. SplatFormer takes as input an initial 3DGS set optimized under limited training views and refines it in a single forward pass, effectively removing potential artifacts in OOD test views. To our knowledge, this is the first successful application of point transformers directly on 3DGS sets, surpassing the limitations of previous multi-scene training methods, which could handle only a restricted number of input views during inference. Our model significantly improves rendering quality under extreme novel views, achieving state-of-the-art performance in these challenging scenarios and outperforming various 3DGS regularization techniques, multi-scene models tailored for sparse view synthesis, and diffusion-based frameworks.

3D高斯点变换器新视角合成渲染优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。