arXiv:2607.18078cs.CV2026-07被引 1

用视觉与几何联合学习,提升纯视觉3D占位预测精度

VGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy Prediction

论文配图:VGOcc: Learning Visual-Geometric Gaussians for Vision-Centric 3D Driving Occupancy Prediction
图 1 · 摘自论文原文
  • 通过视觉-几何高斯初始化与优化,融合多视角深度假设
  • 在nuScenes数据集上达到当前最佳性能,显著提升占位预测精度
  • 适合关注自动驾驶环境感知、3D场景建模的研究者

纯视觉的3D占位预测需从标定的环视图像中恢复语义3D占位场,每张图提供沿相机射线模糊的深度信息。现有方法从密集结构化表示演进到稀疏高斯原语,提升了3D场景表示效率。然而,高斯学习仍主要依赖图像域特征,难以提供显式几何信息以支持体积分析。本文提出VGOcc,从基础模型中学习视觉与几何线索用于高斯建模。VGOcc将这些线索融入原语初始化与精炼过程,形成专为语义占位预测设计的视觉-几何高斯表示。具体而言,提出视觉-几何高斯生成机制,基于射线深度假设生成空间均衡的高斯中心,同时以视觉语义特征初始化原语属性;设计姿态感知特征学习模块,融合基础令牌、相机嵌入与标定射线信息,并在投影3D位置聚合邻近视角特征用于每个高斯精炼阶段;最后,高斯解码器结合姿态感知特征精炼初始高斯并渲染为语义占位。在nuScenes上的实验表明,VGOcc在纯视觉3D占位预测任务中达到当前最优性能。

原文摘要 · Abstract (English)

Vision-only occupancy prediction requires recovering a semantic 3D occupancy field from calibrated surround-view images, where each view provides observations with ambiguous depth along camera rays. Existing methods have progressed from dense structured representations to sparse Gaussian primitives, improving the efficiency of 3D scene representation. However, Gaussian learning still relies primarily on image domain features, which provide limited explicit geometric information for volumetric reasoning. Our key observation is that effective Gaussian occupancy modeling requires not only sparse primitives, but also richer geometric and semantic learning cues. In this paper, we propose VGOcc, which learns visual and geometric cues from foundation models for Gaussian modeling. VGOcc incorporates these cues into primitive initialization and refinement, yielding a representation termed Visual-Geometric Gaussians tailored to semantic occupancy prediction. Specifically, we propose Visual-Geometric Gaussian Birth to form spatially balanced Gaussian centers from ray depth hypotheses, while visual semantic features initialize primitive attributes. Next, we design Pose-Aware Feature Learning to combine foundation tokens with camera embeddings and calibrated ray information. Features from neighboring views are then aggregated at projected 3D locations for each Gaussian refinement stage. Finally, Gaussian decoder refines birth Gaussians with pose-aware features and renders them into semantic occupancy. Experiments on nuScenes demonstrate that VGOcc achieves state-of-the-art performance in vision-only 3D occupancy prediction. Codes will be available at https://github.com/JHLin42in/VGOcc.

3D占位视觉几何高斯建模自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。