重新定义单目3D占据预测评估标准,提升遮挡区域推理性能
Rebenchmarking Unsupervised Monocular 3D Occupancy Prediction
- 基于体素渲染原理重构占据概率表示,实现训练与评估一致
- 引入多视角遮挡感知机制,在遮挡区性能超越现有无监督方法
- 适合关注自动驾驶3D感知、无监督学习的科研与工程人员
从单张图像推断三维结构,特别是在遮挡区域,仍是视觉主导自动驾驶中的基础性难题。现有无监督方法通常训练神经辐射场,并在评估时将网络输出视为占据概率,忽略了训练与评估协议之间的不一致性。此外,普遍采用2D真值无法揭示因几何约束不足导致的遮挡区域固有模糊性。本文提出重构的无监督单目3D占据预测基准。首先解析体素渲染过程中的变量,识别出最物理一致的占据概率表示;基于此改进评估协议,使无监督方法能以与监督方法一致的方式进行评估。为在遮挡区域施加显式约束,引入一种遮挡感知极化机制,融合多视角视觉线索以增强对遮挡区内占据与空闲空间的区分能力。大量实验表明,该方法不仅显著优于现有无监督方法,且达到监督方法的性能水平。源代码与评估协议将在发表后公开。
原文摘要 · Abstract (English)
Inferring the 3D structure from a single image, particularly in occluded regions, remains a fundamental yet unsolved challenge in vision-centric autonomous driving. Existing unsupervised approaches typically train a neural radiance field and treat the network outputs as occupancy probabilities during evaluation, overlooking the inconsistency between training and evaluation protocols. Moreover, the prevalent use of 2D ground truth fails to reveal the inherent ambiguity in occluded regions caused by insufficient geometric constraints. To address these issues, this paper presents a reformulated benchmark for unsupervised monocular 3D occupancy prediction. We first interpret the variables involved in the volume rendering process and identify the most physically consistent representation of the occupancy probability. Building on these analyses, we improve existing evaluation protocols by aligning the newly identified representation with voxel-wise 3D occupancy ground truth, thereby enabling unsupervised methods to be evaluated in a manner consistent with that of supervised approaches. Additionally, to impose explicit constraints in occluded regions, we introduce an occlusion-aware polarization mechanism that incorporates multi-view visual cues to enhance discrimination between occupied and free spaces in these regions. Extensive experiments demonstrate that our approach not only significantly outperforms existing unsupervised approaches but also matches the performance of supervised ones. Our source code and evaluation protocol will be made available upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。