arXiv:2507.19451cs.CV2025-07ICCV被引 16

用视觉数据实现可扩展的3D占位重建,突破了依赖激光雷达的局限。

GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting

  • 基于八叉树高斯表面表示,直接优化显式占位结构。
  • 在Waymo数据集上达到当前最优几何重建效果。
  • 适用于自动驾驶场景的自动标注,适合研究占位重建与视觉感知者。

占位信息对自动驾驶至关重要,为感知与规划提供关键几何先验。然而,现有方法主要依赖激光雷达标注,限制了可扩展性,难以利用大量潜在的众包视觉数据进行自动标注。为此,我们提出GS-Occ3D,一种可扩展的纯视觉占位重建框架,直接重构三维占位。纯视觉方法面临稀疏视角、动态物体、严重遮挡和长时序运动等挑战。现有视觉方法多采用网格表示,存在几何不完整和需额外后处理的问题,影响可扩展性。GS-Occ3D通过八叉树结构的高斯表面表示,优化显式占位,兼顾效率与可扩展性。同时,将场景分解为静态背景、地面和动态物体,分别建模:(1) 地面作为主导结构显式重建,显著提升大范围一致性;(2) 动态车辆独立建模,更好捕捉运动相关占位模式。在Waymo数据集上的实验表明,GS-Occ3D实现了当前最佳的几何重建性能。通过从多样化城市场景中构建纯视觉二值占位标签,验证其在Occ3D-Waymo下游任务中的有效性,以及在Occ3D-nuScenes上的优异零样本泛化能力。这凸显了大规模视觉占位重建作为新型可扩展自动标注范式的潜力。

原文摘要 · Abstract (English)

Occupancy is crucial for autonomous driving, providing essential geometric priors for perception and planning. However, existing methods predominantly rely on LiDAR-based occupancy annotations, which limits scalability and prevents leveraging vast amounts of potential crowdsourced data for auto-labeling. To address this, we propose GS-Occ3D, a scalable vision-only framework that directly reconstructs occupancy. Vision-only occupancy reconstruction poses significant challenges due to sparse viewpoints, dynamic scene elements, severe occlusions, and long-horizon motion. Existing vision-based methods primarily rely on mesh representation, which suffer from incomplete geometry and additional post-processing, limiting scalability. To overcome these issues, GS-Occ3D optimizes an explicit occupancy representation using an Octree-based Gaussian Surfel formulation, ensuring efficiency and scalability. Additionally, we decompose scenes into static background, ground, and dynamic objects, enabling tailored modeling strategies: (1) Ground is explicitly reconstructed as a dominant structural element, significantly improving large-area consistency; (2) Dynamic vehicles are separately modeled to better capture motion-related occupancy patterns. Extensive experiments on the Waymo dataset demonstrate that GS-Occ3D achieves state-of-the-art geometry reconstruction results. By curating vision-only binary occupancy labels from diverse urban scenes, we show their effectiveness for downstream occupancy models on Occ3D-Waymo and superior zero-shot generalization on Occ3D-nuScenes. It highlights the potential of large-scale vision-based occupancy reconstruction as a new paradigm for scalable auto-labeling. Project Page: https://gs-occ3d.github.io/

3D占位视觉重建自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。