无需激光雷达,用视频生成精准3D占位标签,提升视觉占位估计性能。
ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation
- 从视频生成度量一致的3D语义体素标签,实现原生3D监督。
- 在Occ3D-nuScenes上相对之前方法提升34%,超越所有弱监督方法。
- 适合追求无激光雷达3D感知、重视数据质量的研究者。
近期自监督与弱监督占位估计多依赖2D投影或渲染监督,存在几何不一致与深度泄漏问题。本文提出ShelfOcc,一种纯视觉方法,通过视频生成度量一致的语义体素标签,实现无需激光雷达的原生3D监督。尽管当前视觉3D几何基础模型提供先验知识,但其在动态驾驶场景中因几何稀疏、噪声和不一致而难以直接使用。本方法设计专用框架,通过跨帧过滤与累积静态几何,处理动态内容,并将语义信息传播至稳定体素表示。该数据驱动范式使任意先进占位模型架构均可在无激光雷达条件下使用。在Occ3D-nuScenes基准上,ShelfOcc显著优于所有先前弱/货架监督方法(相对提升最高达34%),为无激光雷达3D场景理解开辟新路径。
原文摘要 · Abstract (English)
Recent progress in self- and weakly supervised occupancy estimation has largely relied on 2D projection or rendering-based supervision, which suffers from geometric inconsistencies and severe depth bleeding. We thus introduce ShelfOcc, a vision-only method that overcomes these limitations without relying on LiDAR. ShelfOcc brings supervision into native 3D space by generating metrically consistent semantic voxel labels from video, enabling true 3D supervision without any additional sensors or manual 3D annotations. While recent vision-based 3D geometry foundation models provide a promising source of prior knowledge, they do not work out of the box as a prediction due to sparse or noisy and inconsistent geometry, especially in dynamic driving scenes. Our method introduces a dedicated framework that mitigates these issues by filtering and accumulating static geometry consistently across frames, handling dynamic content and propagating semantic information into a stable voxel representation. This data-centric shift in supervision for weakly/shelf-supervised occupancy estimation allows the use of essentially any SOTA occupancy model architecture without relying on LiDAR data. We argue that such high-quality supervision is essential for robust occupancy learning and constitutes an important complementary avenue to architectural innovation. On the Occ3D-nuScenes benchmark, ShelfOcc substantially outperforms all previous weakly/shelf-supervised methods (up to a 34% relative improvement), establishing a new data-driven direction for LiDAR-free 3D scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。