arXiv:2603.06210cs.CVcs.RO2026-03中稿 · IROS 2026被引 4

用视觉大模型增强3D高斯点云,提升自动驾驶语义占位预测精度

VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction

  • 引入视觉大模型几何先验,通过分层特征适配器融合多视角信息
  • 在nuScenes数据集上提升12.6%的IoU和7.5%的mIoU
  • 可通用适配多种视觉大模型,适合追求高精度感知的自动驾驶研究者

3D语义占位预测已成为自动驾驶中全面场景理解的关键感知任务。尽管近期研究利用3D高斯点云建模显著降低了计算开销,但高质量3D高斯点的生成仍严重依赖精确的几何线索,而纯视觉范式常缺乏此类信息。为此,我们提出将视觉基础模型(VFMs)强大的几何先验注入占位预测。本文提出视觉几何接地高斯点云(VG3S),一种基于高斯的占位预测新框架,实现跨视角3D几何对齐。具体地,为充分挖掘冻结的VFM中丰富的3D几何先验,设计了一种即插即用的分层几何特征适配器,通过特征聚合、任务对齐与多尺度重构,有效转换通用的VFM token。在nuScenes占位基准上的大量实验表明,相比基线,VG3S在IoU上提升12.6%,在mIoU上提升7.5%。此外,验证了其在多种不同VFMs间的良好泛化能力,持续提升占位预测精度,凸显了集成预训练几何引导型大模型先验的巨大价值。

原文摘要 · Abstract (English)

3D semantic occupancy prediction has become a crucial perception task for comprehensive scene understanding in autonomous driving. While recent advances have explored 3D Gaussian splatting for occupancy modeling to substantially reduce computational overhead, the generation of high-quality 3D Gaussians relies heavily on accurate geometric cues, which are often insufficient in purely vision-centric paradigms. To bridge this gap, we advocate for injecting the strong geometric grounding capability from Vision Foundation Models (VFMs) into occupancy prediction. In this regard, we introduce Visual Geometry Grounded Gaussian Splatting (VG3S), a novel framework that empowers Gaussian-based occupancy prediction with cross-view 3D geometric grounding. Specifically, to fully exploit the rich 3D geometric priors from a frozen VFM, we propose a plug-and-play hierarchical geometric feature adapter, which can effectively transform generic VFM tokens via feature aggregation, task-specific alignment, and multi-scale restructuring. Extensive experiments on the nuScenes occupancy benchmark demonstrate that VG3S achieves remarkable improvements of 12.6% in IoU and 7.5% in mIoU over the baseline. Furthermore, we show that VG3S generalizes seamlessly across diverse VFMs, consistently enhancing occupancy prediction accuracy and firmly underscoring the immense value of integrating priors derived from powerful, pre-trained geometry-grounded VFMs.

3D感知高斯点云语义占位视觉大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。