用体素高斯点云实现无需优化的3D重建,效果超越现有方法。
$\text{VG}^2$GT: Voxel-Gaussian Splatting Visual Geometry Grounded Transformer

- 通过体素特征直接生成高斯参数,避免传统优化流程。
- 在DTU、Replica等数据集上显著优于当前最优方法。
- 可适配任意基于图像块的视觉模型,训练成本大幅降低。
高斯点阵在3D重建和新视角合成中展现出强大潜力。然而,现有方法通常需要精确相机参数和场景级优化,而采用像素对齐高斯原型的前馈方法常出现伪影且分布不均。本文提出VG²GT:一种体素-高斯点阵视觉几何接地变换器。VG²GT利用冻结的预训练视觉基础模型(VFM),引入多尺度可微体素模块以增强几何理解,并直接从体素特征中分裂与回归高斯原型参数。训练时通过随机体积极体渲染监督深度图,实现几何准确的高斯场景重建,同时保持视觉基础模型完全冻结。该设计使VG²GT可无缝嵌入任意基于图像块的VFM,显著降低训练成本。在广泛使用的DTU、Replica、TAT和ScanNet数据集上,其性能超越当前最先进方法。
原文摘要 · Abstract (English)
Gaussian splatting has shown strong potential for 3D reconstruction and novel view synthesis. However, most existing methods require accurate camera parameters and per-scene optimization, while feed-forward methods with pixel-aligned Gaussian primitives often suffer from artifacts and non-uniform primitives. In this paper, we propose $\text{VG}^2$GT, a Voxel-Gaussian Splatting Visual Geometry-Grounded Transformer. $\text{VG}^2$GT leverages a frozen pretrained visual foundation model (VFM), incorporates a multi-scale differentiable voxel module to enhance geometric understanding, and directly splits and regresses Gaussian primitive parameters from voxel features. During training, depth maps are supervised through stochastic solid volume rendering, enabling geometrically accurate Gaussian scene reconstruction while keeping the visual foundation model fully frozen. This design enables $\text{VG}^2$GT to be seamlessly plugged into any patch-feature-based VFM, while substantially reducing the required training cost. $\text{VG}^2$GT outperforms current state-of-the-art methods on widely used DTU, Replica, TAT, and ScanNet datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。