arXiv:2510.19578cs.CV2025-10被引 8

用视觉几何先验提升自动驾驶全景重建质量

VGD: Visual Geometry Gaussian Splatting for Feed-Forward Surround-view Driving Reconstruction

  • 通过轻量VGGT提取几何先验,显式建模场景结构
  • 多尺度高斯头融合几何信息,实现高质量新视角渲染
  • 适合追求高保真全景重建的自动驾驶研究者

前向式全景自动驾驶场景重建具有快速、泛化性强的优势,但如何在保证泛化性的同时提升新视角质量仍是核心挑战。由于全景视角重叠区域极少,现有方法难以保障新视角的几何一致性与重建质量。为此,本文提出视觉高斯驾驶(VGD)框架,显式学习几何信息并用于引导语义质量提升。设计轻量版VGGT架构,从预训练模型中高效提取几何先验至几何分支;引入高斯头,融合多尺度几何特征以预测新视角渲染所需高斯参数,共享同一图像块骨干网络;最后将几何与高斯分支的多尺度特征联合监督语义优化模型,通过特征一致性学习提升渲染质量。在nuScenes数据集上的实验表明,本方法在多种设置下均显著优于当前最优方法,在客观指标与主观质量上均有提升,验证了其可扩展性与高保真全景重建能力。

原文摘要 · Abstract (English)

Feed-forward surround-view autonomous driving scene reconstruction offers fast, generalizable inference ability, which faces the core challenge of ensuring generalization while elevating novel view quality. Due to the surround-view with minimal overlap regions, existing methods typically fail to ensure geometric consistency and reconstruction quality for novel views. To tackle this tension, we claim that geometric information must be learned explicitly, and the resulting features should be leveraged to guide the elevating of semantic quality in novel views. In this paper, we introduce \textbf{Visual Gaussian Driving (VGD)}, a novel feed-forward end-to-end learning framework designed to address this challenge. To achieve generalizable geometric estimation, we design a lightweight variant of the VGGT architecture to efficiently distill its geometric priors from the pre-trained VGGT to the geometry branch. Furthermore, we design a Gaussian Head that fuses multi-scale geometry tokens to predict Gaussian parameters for novel view rendering, which shares the same patch backbone as the geometry branch. Finally, we integrate multi-scale features from both geometry and Gaussian head branches to jointly supervise a semantic refinement model, optimizing rendering quality through feature-consistent learning. Experiments on nuScenes demonstrate that our approach significantly outperforms state-of-the-art methods in both objective metrics and subjective quality under various settings, which validates VGD's scalability and high-fidelity surround-view reconstruction.

自动驾驶全景重建高斯溅射几何先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。