arXiv:2608.10682cs.CV2026-08

用视觉几何先验提升单帧环视重建的几何稳定性。

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

论文配图:Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction
图 1 · 摘自论文原文
  • 引入VGGT提供可迁移的多视角几何先验
  • 在nuScenes上实现最佳渲染质量与几何一致性
  • 适合自动驾驶环视场景的快速重建应用

单帧环视重建因相机间重叠极少而面临严重几何不稳和渲染伪影。现有方法依赖复杂解码器或辅助信号,但受限于上游特征的弱几何表达能力。本文提出VGGD,一种面向驾驶场景的视觉几何基础感知3D高斯点阵框架,将几何建模前置并适配预训练先验。首先,利用VGGT生成可迁移的多视图几何先验令牌;其次,设计双路径颈部模块分离几何一致与外观感知表示,提升弱观测区域的外观补全;再通过尺度热身策略稳定早期几何学习,抑制自身位姿变化下的尺度漂移;最后采用像素-体素混合高斯解码器生成可渲染3D高斯场景用于新视角合成。在nuScenes单帧基准测试中,VGGD在所有对比方法中取得最优整体渲染质量,并显著提升相对几何一致性。

原文摘要 · Abstract (English)

Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing methods rely on complex decoders or auxiliary cues, they remain bottlenecked by the weak geometric capacity of upstream features. We argue that leveraging pretrained visual geometry priors strengthens upstream representations and alleviates the geometric ambiguity in sparse surround views. To this end, we propose VGGD, a visual geometry foundation-aware 3D Gaussian Splatting framework for feed-forward surround-view driving reconstruction, which shifts geometric modeling to the frontend and adapts foundation priors to the driving camera setting. First, VGGD leverages VGGT to provide transferable multi-view geometric prior tokens. Next, we introduce a Dual-Path Neck to decouple geometry-consistent and appearance-aware representations, improving appearance completion in weakly observed regions. We further apply Scale Warmup to stabilize early geometry learning and suppress scale drift under ego-pose changes. Finally, we use a hybrid pixel--volume Gaussian decoder to produce a renderable 3D Gaussian scene for novel-view synthesis. Experiments on the nuScenes single-frame benchmark show that VGGD achieves the best overall rendering quality among the compared methods and improves relative geometric consistency.

3D重建高斯点阵自动驾驶几何先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。