arXiv:2607.13481cs.CV2026-07

用视觉几何先验构建稀疏高斯占据表示,实现高效3D场景理解。

GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors

论文配图:GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors
图 1 · 摘自论文原文
  • 将视觉几何先验转化为体积化的稀疏高斯表示
  • 在多视角与时序场景中均达到强性能,效率与泛化性俱佳
  • 适用于室内外复杂场景,适合自动驾驶与机器人感知

精确的3D场景理解是具身智能与自动驾驶的基础,其中3D占据表征统一表达了物体、结构与自由空间。然而,从视觉观测中恢复完整的体素表示仍具挑战性,尤其在遮挡和未观测区域。视觉几何先验提供了强大且可泛化的几何线索,但其输出本质为表面中心,而占据预测需推理体内部与自由空间。为此,我们提出GPOcc,将视觉几何先验转换为占据感知的稀疏高斯表示,实现高效且表达丰富的体素场景建模。基于GPOcc,GPOcc++在统一框架内建模多视角观测与时间序列,使空间与时间证据通过同一表示处理。我们进一步将GPOcc++从室内场景扩展至室外占据预测。在室内外基准上的大量实验表明,其在多视角与时序设置下均表现出一致强性能,兼具良好效率与泛化能力。代码将发布于https://github.com/JuIvyy/GPOcc。

原文摘要 · Abstract (English)

Accurate 3D scene understanding is fundamental to embodied intelligence and autonomous driving, where 3D occupancy provides a unified representation of objects, structures, and free space. However, recovering such a complete volumetric representation from visual observations remains challenging, particularly in occluded and unobserved regions. Visual geometry priors offer strong and generalizable geometric cues for addressing this challenge, but their outputs are inherently surface-centric, whereas occupancy prediction requires reasoning about volumetric interiors and free space. To bridge this gap, we introduce GPOcc, which transforms visual geometry priors into occupancy-aware sparse Gaussian representations for efficient and expressive volumetric scene modeling. Building on GPOcc, GPOcc++ models multi-view observations and temporal sequences within a unified framework, allowing spatial and temporal evidence to be handled through the same representation. We further extend GPOcc++ from indoor scenes to outdoor occupancy prediction. Extensive experiments on both indoor and outdoor benchmarks demonstrate consistently strong performance across both multi-view and temporal settings, together with favorable efficiency and generalization. Code will be released at https://github.com/JuIvyy/GPOcc.

3D占据视觉几何稀疏表示自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。