arXiv:2609.04846cs.CV2026-09

用多视角投票提升2D伪标签可靠性,实现无需3D标注的3D占据预测

LetOccVote: Learning Weakly Supervised 3D Occupancy through Consensus

论文配图:LetOccVote: Learning Weakly Supervised 3D Occupancy through Consensus
图 1 · 摘自论文原文
  • 通过跨帧几何与语义投票,筛选可信2D伪标签
  • 在Occ3D-nuScenes上达到53.27的IoU和20.39的mIoU
  • 适合无3D标注但需高精度3D占据建模的场景

弱监督3D占据预测通过视觉基础模型生成的2D伪标签降低对昂贵3D标注的依赖。然而,现有方法直接使用不完美的伪标签作为监督信号,使占据学习易受错误几何和语义目标影响。我们观察到:重复观测间的共识是评估伪标签可靠性的低成本且可靠的线索。基于此,提出 extbf{LetOccVote}——一种基于高斯的弱监督占据框架,利用跨帧投票增强几何与语义监督。几何方面,深度投票通过跨帧几何一致性精炼伪深度并剔除矛盾估计,再进行体素提升与深度监督;语义方面,语义投票在共享3D空间聚合伪语义观测,识别可靠与争议证据,强化可信语义监督并过滤不可靠伪标签段。整个框架仅需2D伪标签监督,无需3D占据标注。在Occ3D-nuScenes上,取得53.27 IoU与20.39 mIoU,超越现有2D伪标签方法的最先进水平。

原文摘要 · Abstract (English)

Weakly supervised 3D occupancy prediction reduces the reliance on costly 3D annotations by learning from 2D pseudo-labels generated by vision foundation models. However, existing methods typically use these imperfect pseudo-labels directly as supervision, making occupancy learning vulnerable to erroneous geometric and semantic targets. We observe that agreement across repeated observations provides an inexpensive and reliable cue for assessing pseudo-label reliability. Based on this observation, we propose \textbf{LetOccVote}, a weakly supervised Gaussian-based occupancy framework that leverages cross-frame voting to improve both geometric and semantic supervision. For geometry, Depth Vote exploits cross-frame geometric agreement to refine supported pseudo depth and reject contradictory estimates before volumetric lifting and depth supervision. For semantics, Semantic Vote aggregates pseudo-semantic observations in a shared 3D space to identify reliable and contested evidence, strengthening reliable semantic supervision while filtering unreliable pseudo-label segments. The entire framework is trained solely with 2D pseudo-label supervision without requiring 3D occupancy annotations. On Occ3D-nuScenes, LetOccVote achieves 53.27 IoU and 20.39 mIoU, establishing state-of-the-art performance among methods with 2D pseudo-label supervision.

3D占据弱监督伪标签多视角投票

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。